@relaygpu/mcp
Relay's MCP server. Your agent can find a Relay model, read its input schema, run it, poll long tasks, upload files, run workflows and search the Relay docs, all from Claude Code, Cursor, VS Code or Claude Desktop. It is built on @relaygpu/client. The tools are generic and take the model as a string, so a model added to Relay works without a new release.
Run it locally; your client starts it on demand:
| How | Credential | |
|---|---|---|
| Local (stdio) | npx -y @relaygpu/mcp (Node ≥ 20) |
RELAY_API_KEY=relay_sk_... in its environment; optional RELAY_BASE_URL (default https://relaygpu.com) |
| Hosted (Streamable HTTP) — coming soon | https://mcp.relaygpu.com/mcp (not live yet) |
X-API-Key: relay_sk_... header, or Authorization: Bearer <relay key or dashboard JWT> |
The server never stores your key: the local server reads it from its environment, the hosted one will read it from each request and use it for that request only.
Install
Every client below runs the same local server with npx; put your Relay key where the snippet says. Hosted connections (OAuth, Claude Desktop's connector directory, Codex remote servers) come with the hosted endpoint.
Claude Code
claude mcp add relay --env RELAY_API_KEY=relay_sk_... -- npx -y @relaygpu/mcp
Codex
codex mcp add relay --env RELAY_API_KEY=relay_sk_... -- npx -y @relaygpu/mcp
Or in ~/.codex/config.toml (project-level .codex/config.toml works for trusted projects):
[mcp_servers.relay]
command = "npx"
args = ["-y", "@relaygpu/mcp"]
env = { RELAY_API_KEY = "relay_sk_..." }
Cursor
The button installs the entry below with a placeholder key; replace relay_sk_... in Cursor's MCP settings afterwards. By hand: ~/.cursor/mcp.json (all projects) or .cursor/mcp.json (one project).
{
"mcpServers": {
"relay": {
"command": "npx",
"args": ["-y", "@relaygpu/mcp"],
"env": { "RELAY_API_KEY": "relay_sk_..." }
}
}
}
VS Code
VS Code asks for the key once and keeps it in its secret storage. By hand, in .vscode/mcp.json:
{
"inputs": [{ "type": "promptString", "id": "relay-key", "description": "Relay API key (relay_sk_...)", "password": true }],
"servers": {
"relay": {
"type": "stdio",
"command": "npx",
"args": ["-y", "@relaygpu/mcp"],
"env": { "RELAY_API_KEY": "${input:relay-key}" }
}
}
}
Claude Desktop (stdio only)
Claude Desktop, claude.ai connectors and ChatGPT accept remote MCP servers only through OAuth, and v0 does not offer OAuth, so these clients use the local server. Add the following to claude_desktop_config.json (Settings → Developer → Edit Config), then restart Claude Desktop:
{
"mcpServers": {
"relay": {
"command": "npx",
"args": ["-y", "@relaygpu/mcp"],
"env": { "RELAY_API_KEY": "relay_sk_..." }
}
}
}
Ready-to-copy files are in examples/.
Tools
Tools marked keyless work without a credential. The others are billed or account tools and need one.
| Tool | What it does | Example |
|---|---|---|
search_models keyless |
Find models by tag and/or text | "Find a text-to-video model from Kling" → {tag: "text-to-video", query: "kling"} |
get_model keyless |
Route, request schema, example, pricing and status of one model | {model: "Qwen/qwen-image"} |
get_pricing keyless |
Public prices for one model or all of them, including the media_storage SKUs |
{model: "KlingTeam/v3-T2V"} |
estimate_cost keyless |
Estimate from the public price list (not an invoice) | {model: "KlingTeam/v3-T2V", usage: {duration_seconds: 5, quality_mode: "std"}} |
run_model |
Run any model by name and wait up to 45 s | {model: "Qwen/qwen-image", input: {prompt: "a red fox", size: "512x512"}} |
check_task keyless |
Poll a task (waits up to 30 s) | {task_id: "direct:…", wait_seconds: 30} |
upload_file |
Host a file and get a URL for any *_url input |
{path: "./clip.mp4", retention: "relay1d"} |
list_workflows |
List the workflow templates | "Which workflows can I run?" |
run_workflow |
Start a workflow run, wait up to 45 s | {workflow_id: "script-voiceover", inputs: {messages: [{role: "user", content: "a lighthouse at dusk"}], voice: "Cherry"}} |
check_workflow_run keyless |
Poll a run (waits up to 30 s) | {run_id: "…", wait_seconds: 30} |
cancel_workflow_run |
Cancel a run that is still in progress | {run_id: "…"} |
get_credits |
Credit balance, promos and consumption (JWT or custom-tier superkey) | "How many credits do I have left?" |
get_usage |
Usage analytics (JWT or custom-tier superkey) | {period: "7d"} |
search_docs keyless |
Search the Relay docs | {query: "idempotency key"} |
Notes
- Auth.
tools/listand the keyless tools answer without a credential. A billed tool called without one returns aMISSING_CREDENTIALStool error that names what to set: theX-API-Keyheader (hosted) orRELAY_API_KEY(stdio). Errors from Relay come back as tool errors that include the error class,code,requestIdand the server's detail. Your key never appears in a tool result or a log line. - Long tasks.
run_modelwaits up to 45 s (wait_seconds, 0–45). If the task is still running after that, it returns thetask_idand the line "Do not resubmit: call check_task with this task_id". Every submit carries anIdempotency-Key, so a retry does not bill twice. Pass your ownidempotency_keyif your agent might send the same submit again. - Expiry. Each result link comes with its expiry: "expires 1 h after completion unless store_output was set". To keep results longer, pass
store_outputwith amedia_storageSKU fromget_pricing(for examplerelay7d). That SKU adds a per-file fee. - Re-hosting. Some routes return images only as inline base64 (for example gpt-image and Flux). The server uploads those images through
/v2/files, at thestore_outputretention orrelay1hby default (free within the daily quota), and returns a URL and afile_idinstead of the base64. Image content blocks are opt-in: passinline_images: true, which applies to images of 1 MB or less. - Uploads. On stdio,
upload_filestreams a localpathof up to 100 MB. The hosted server cannot read your disk. On hosted, passbase64(up to 4 MB) or a publicurlinstead. Retention defaults torelay1h.
Self-hosting
docker build -t relaygpu-mcp .
docker run -p 2300:2300 -e PORT=2300 -e RELAY_BASE_URL=https://relaygpu.com relaygpu-mcp
The server listens on PORT (default 2300) and calls Relay at RELAY_BASE_URL. Endpoints:
POST /mcp: Streamable HTTP, stateless.GET /healthz: returns{status, version}.GET /metrics: Prometheus text withmcp_tool_calls_total{tool,outcome}andmcp_tool_seconds{tool}.GET /.well-known/mcp/server-card.json: the server and its tools.
Logs are JSON lines on stderr, WARNING and above by default (LOG_LEVEL). Without Docker, run npm ci && npm run build && node dist/http.js.
License
MIT