Claude Code
Run Claude Code on deeprelay models: one setup command points every model slot at the Anthropic Messages endpoint, with what works, what is rejected, and how to undo it.
deeprelay serves the Anthropic Messages API (POST /v1/messages), so
Claude Code can run on
deeprelay models. Claude Code keeps working as usual: it reads and edits
files, runs commands and calls tools. Every request goes to a deeprelay model
you picked, and it is billed like any other deeprelay inference.
deeprelay does not serve Claude models. Setup points every model slot Claude Code uses at a deeprelay model, so a configured session never asks for one.
Quick start
deeprelay login # once: stores your key in the deeprelay credentials file
deeprelay setup claude-code # points Claude Code at deeprelay (asks before writing)
claude # start Claude Code as usual
Setup checks your key and every chosen model before it writes anything. It then shows a diff of the settings file and the undo command, and asks you to confirm.
What setup changes
deeprelay setup claude-code merges these keys into a Claude Code settings
file.
| Key | Value |
|---|---|
env.ANTHROPIC_BASE_URL | https://api.deeprelay.ai (no /v1: Claude Code adds /v1/messages itself) |
env.ANTHROPIC_MODEL | the main model (--model) |
env.ANTHROPIC_DEFAULT_OPUS_MODEL | the model behind the opus alias (--opus-model) |
env.ANTHROPIC_DEFAULT_SONNET_MODEL | the model behind the sonnet alias (--sonnet-model) |
env.ANTHROPIC_DEFAULT_HAIKU_MODEL | the model behind the haiku alias (--haiku-model) |
env.ANTHROPIC_DEFAULT_FABLE_MODEL | the model behind the fable alias (--fable-model) |
env.ANTHROPIC_SMALL_FAST_MODEL | the background-task model (same as --haiku-model) |
apiKeyHelper | the absolute path of your deeprelay binary, followed by auth token |
Setup also removes env.ANTHROPIC_AUTH_TOKEN and env.ANTHROPIC_API_KEY
from the file if they are there, and turns off a cloud-provider mode
(CLAUDE_CODE_USE_BEDROCK, _VERTEX or _FOUNDRY) that is switched on.
Claude Code uses any of these in preference to apiKeyHelper, so a key left
over from another gateway would be sent to deeprelay instead of yours. The
diff shows each removal; a credential's value is hidden, while a cloud-mode
flag shows its true/false value. --remove puts the old values back.
Everything else in the file is left alone.
Your key is never written to the settings file. apiKeyHelper makes
Claude Code run deeprelay auth token whenever it needs the key, so the key
stays in the deeprelay credentials file. The binary is referenced by its
absolute path. If you move or reinstall deeprelay, run setup again.
Where it writes
- User scope (default):
~/.claude/settings.json, or$CLAUDE_CONFIG_DIR/settings.jsonif you set that variable. Every project on this machine uses deeprelay. - Project scope (
--scope project):.claude/settings.local.jsonin the current directory. Only this project uses deeprelay. Setup warns if git does not ignore that file.
Backups and undo
Before each write, setup copies the current settings file into the
backups/claude-code folder of the deeprelay config directory
(~/.config/deeprelay, or $XDG_CONFIG_HOME/deeprelay). Backups are
timestamped and never overwritten. Setup also records which keys it added,
changed or removed, and their earlier values. Your deeprelay key is never in
that record; a credential setup removed from the settings file is, so that
--remove can restore it.
deeprelay setup claude-code --remove replays that record. Add
--scope project to undo a project setup.
- Keys setup added are deleted.
- Keys setup changed get their earlier values back.
- Keys setup removed (a conflicting credential) are put back.
- Keys you edited yourself after setup are left alone, with a warning.
Warnings
These always print, even with --yes:
- the settings file already points
ANTHROPIC_BASE_URLsomewhere else, so setup would redirect all Claude Code usage to deeprelay; - the settings file holds a credential or cloud-provider setting that outranks
apiKeyHelper, which setup removes (see above); - with
--scope project, your user settings file holds such a setting; setup does not edit that file, and Claude Code will not reach deeprelay until you remove it there; ANTHROPIC_*variables are set in your shell, which override the settings file (only the names are shown);- Claude Code is logged in to a Claude subscription and you chose user scope, so every project routes to deeprelay instead. Consider
--scope project.
--print for shells and CI
deeprelay setup claude-code --print writes nothing. It prints export
lines for the same variables, with ANTHROPIC_AUTH_TOKEN in place of
apiKeyHelper. That output contains your key in plaintext. Don't paste
it into a file you commit or a log you share. In CI, prefer a secret store
that sets ANTHROPIC_AUTH_TOKEN for the job.
Models
Setup's defaults are models that completed real Claude Code coding sessions
(reading files, editing, running commands, thinking, and a follow-up turn
with --continue):
| Slot | Default | Why |
|---|---|---|
Main, Opus, Sonnet, Fable (--model, --opus-model, --sonnet-model, --fable-model) | deeprelay/deepseek-v4-pro | The strongest tested model with thinking and reliable tool calls |
Haiku and small-fast (--haiku-model) | deeprelay/deepseek-v4.1-flash | Covered by the flat plan (a subscriber's background tasks stay inside it), the same family as the main model, and it can see images; used for background tasks such as titles and summaries |
deeprelay/glm-5.3 also completed a full session and works as a main model:
deeprelay setup claude-code --model deeprelay/glm-5.3
deeprelay/gpt-oss-120b completed a full session as well. It is not covered
by the flat plan, but it is a cheaper pay-as-you-go choice for the background
slot: --haiku-model deeprelay/gpt-oss-120b.
Any other model works if it supports tool calling. Setup checks this and refuses a model without it. To see what is available right now:
deeprelay models list
The catalog is live, so a model can come and go. If a model stops being served, run setup again with another one.
Why claude-* model names are rejected
A request for a Claude model ID (for example claude-sonnet-…) gets a 400
instead of being quietly mapped to some other model. deeprelay doesn't alias
Claude names, so you always know which model answered, and you are billed
for the model you chose. If you see this error, a model slot is not set: run
deeprelay setup claude-code, or set the variables below.
Manual setup
If you'd rather not use deeprelay setup, set the same variables yourself,
in your shell or in the env block of a Claude Code settings file:
export ANTHROPIC_BASE_URL=https://api.deeprelay.ai
export ANTHROPIC_AUTH_TOKEN=deeprelay_live_...
export ANTHROPIC_MODEL=deeprelay/deepseek-v4-pro
export ANTHROPIC_DEFAULT_OPUS_MODEL=deeprelay/deepseek-v4-pro
export ANTHROPIC_DEFAULT_SONNET_MODEL=deeprelay/deepseek-v4-pro
export ANTHROPIC_DEFAULT_FABLE_MODEL=deeprelay/deepseek-v4-pro
export ANTHROPIC_DEFAULT_HAIKU_MODEL=deeprelay/deepseek-v4.1-flash
export ANTHROPIC_SMALL_FAST_MODEL=deeprelay/deepseek-v4.1-flash
ANTHROPIC_BASE_URLhas no/v1. Claude Code appends/v1/messagesitself, sohttps://api.deeprelay.ai/v1would end up at/v1/v1/messages.- Set every model variable. A slot you leave unset falls back to a Claude model name, and that request is rejected.
- Instead of
ANTHROPIC_AUTH_TOKEN, you can put"apiKeyHelper": "/absolute/path/to/deeprelay auth token"in the settings file, so the key never sits in your shell or a file. - deeprelay accepts the key as
x-api-keyor asAuthorization: Bearer, soANTHROPIC_API_KEYworks too.
What works and what doesn't
Works:
- Tools. File reads and edits, shell commands, subagents and MCP servers you configure in Claude Code all run through normal tool calls.
- Streaming, including keep-alive pings during long responses.
- Thinking. Models that reason stream it as thinking blocks. deeprelay signs them, and handles them correctly when Claude Code sends them back on later turns.
- Prompt caching. Caching happens automatically on models that support it.
cache_controlmarkers are accepted, and cached input shows up ascache_read_input_tokensin usage.
Doesn't work:
- WebSearch. Claude Code's built-in WebSearch tool relies on an Anthropic server-side tool, which deeprelay does not run. The request returns a
400that says so, Claude Code shows the error to the model, and the session carries on. For web search, add a search MCP server to Claude Code instead. The same applies to other server tools, such as code execution. WebFetch runs on your machine as an ordinary client tool, so it is not affected. - PDFs and documents. Document blocks are rejected with a
400. - Images, except on vision models. Today only
deeprelay/deepseek-v4.1-flashcan see images. On that model, images you paste and images Claude Code's Read tool opens (a screenshot or a PNG in your repo) reach the model. On other models, images never break the session:- an image Read opens, or an image from an earlier turn, is replaced by a note telling the model that an image was there and it cannot see it, and the session carries on;
- only an image you paste into your newest message returns a
400, which names the models that can see images. Remove the image or switch models (/model).
When deeprelay replaces or moves an image, the response carries the header
x-deeprelay-ignored-params with image:omitted (replaced by a note) or
image:moved (on a vision model, an image from a tool result is passed to
the model right after the tool result, because most models accept tool
results as text only). Forwarded images must be JPEG, PNG, GIF or WebP, at
most 5 MiB each and 20 per request.
- Exact token counts.
POST /v1/messages/count_tokensreturns an estimate, marked with the response headerx-deeprelay-token-count: estimated. It tends to count high, which is safe for Claude Code's context budgeting. It is free.
Billing
- A Claude Code request is billed exactly like a chat completion for the same model: the same per-token prices, and the same plan quota and credit.
- Rejected requests are not billed. That includes a
claude-*model name, a web search call and an unsupported image. Token counting is never billed. - Usage and spend show the deeprelay model that answered. Check them with
deeprelay usageor in the dashboard. - Claude Code's
/costdoes not show deeprelay prices. It computes cost from Anthropic's price list, which doesn't apply to deeprelay models. For what you actually spent, usedeeprelay usageor the usage page in the deeprelay dashboard.
Troubleshooting
Errors on /v1/messages use the Anthropic error format, so Claude Code shows
them as readable messages. Each error also carries an
X-Deeprelay-Error-Code header with the same code the rest of the API uses
(see Errors).
"deeprelay doesn't serve Claude models"
deeprelay doesn't serve Claude models — run `deeprelay setup claude-code`, or set ANTHROPIC_MODEL, ANTHROPIC_DEFAULT_{OPUS,SONNET,HAIKU,FABLE}_MODEL and ANTHROPIC_SMALL_FAST_MODEL to deeprelay model IDs (see deeprelay.ai/docs/claude-code).
Claude Code asked for a Claude model, so at least one model slot is not set
(code claude_model_not_served). Run deeprelay setup claude-code again, or
set every variable listed under Manual setup, including
ANTHROPIC_DEFAULT_FABLE_MODEL and ANTHROPIC_SMALL_FAST_MODEL.
"apiKeyHelper is failing"
Claude Code couldn't get a key from the helper. Run the helper yourself to see why:
deeprelay auth token
Usually you are not logged in (run deeprelay login), or the deeprelay
binary moved since setup ran (run setup again).
Settings seem to be ignored
- Shell variables win. An
ANTHROPIC_*variable exported in your shell overrides the settings file. Setup lists any it finds. Unset them, or use a clean shell. - Claude Code is still using your subscription. If you are logged in to a Claude subscription and ran setup at project scope, only that project uses deeprelay. Check which settings file setup wrote, and run
/statusin Claude Code to see the active base URL and model.
"prompt is too long"
The conversation no longer fits the model's context window (code
context_length_exceeded). deeprelay returns Anthropic's own
prompt is too long: N tokens > M maximum error, so Claude Code handles it
as it would with Anthropic: it compacts the conversation, or asks you to run
/compact. The token count is deeprelay's estimate, and M is the model's
context window from deeprelay models get.
402 and 429
| Status | X-Deeprelay-Error-Code | Meaning |
|---|---|---|
401 | unauthorized | The key is missing, invalid, revoked or expired |
402 | insufficient_balance | Out of credit. Top up in the dashboard |
402 | daily_limit_reached, spending_limit_reached | A spend limit on your account or key was hit |
429 | subscription_quota_exhausted | Your plan's quota for this window is used up |
429 | rate_limit_exceeded | Too many requests per second. Claude Code retries |
429 | stream_limit_exceeded | Too many requests open at once on this key. Claude Code retries |
Billing 402s and an exhausted plan quota tell Claude Code not to retry, so
it reports them once instead of looping. Check where you stand with
deeprelay billing subscription and deeprelay usage.
Still stuck? Email support@deeprelay.ai with your Claude Code version, the
model ID and the exact error message.
Related
- Coding agents: Cline, Kilo Code, Roo Code, OpenCode, Aider, Continue, Zed and Cursor
- Serverless inference API: the
/v1/messagesendpoint in detail deeprelay setup claude-code: every flagdeeprelay auth token: the key helper