deeprelayDocs
CH·GGuides

Claude Code

Run Claude Code on deeprelay models: one setup command points every model slot at the Anthropic Messages endpoint, with what works, what is rejected, and how to undo it.

deeprelay serves the Anthropic Messages API (POST /v1/messages), so Claude Code can run on deeprelay models. Claude Code keeps working as usual: it reads and edits files, runs commands and calls tools. Every request goes to a deeprelay model you picked, and it is billed like any other deeprelay inference.

deeprelay does not serve Claude models. Setup points every model slot Claude Code uses at a deeprelay model, so a configured session never asks for one.

Quick start

deeprelay login                 # once: stores your key in the deeprelay credentials file
deeprelay setup claude-code     # points Claude Code at deeprelay (asks before writing)
claude                          # start Claude Code as usual

Setup checks your key and every chosen model before it writes anything. It then shows a diff of the settings file and the undo command, and asks you to confirm.

What setup changes

deeprelay setup claude-code merges these keys into a Claude Code settings file.

KeyValue
env.ANTHROPIC_BASE_URLhttps://api.deeprelay.ai (no /v1: Claude Code adds /v1/messages itself)
env.ANTHROPIC_MODELthe main model (--model)
env.ANTHROPIC_DEFAULT_OPUS_MODELthe model behind the opus alias (--opus-model)
env.ANTHROPIC_DEFAULT_SONNET_MODELthe model behind the sonnet alias (--sonnet-model)
env.ANTHROPIC_DEFAULT_HAIKU_MODELthe model behind the haiku alias (--haiku-model)
env.ANTHROPIC_DEFAULT_FABLE_MODELthe model behind the fable alias (--fable-model)
env.ANTHROPIC_SMALL_FAST_MODELthe background-task model (same as --haiku-model)
apiKeyHelperthe absolute path of your deeprelay binary, followed by auth token

Setup also removes env.ANTHROPIC_AUTH_TOKEN and env.ANTHROPIC_API_KEY from the file if they are there, and turns off a cloud-provider mode (CLAUDE_CODE_USE_BEDROCK, _VERTEX or _FOUNDRY) that is switched on. Claude Code uses any of these in preference to apiKeyHelper, so a key left over from another gateway would be sent to deeprelay instead of yours. The diff shows each removal; a credential's value is hidden, while a cloud-mode flag shows its true/false value. --remove puts the old values back. Everything else in the file is left alone.

Your key is never written to the settings file. apiKeyHelper makes Claude Code run deeprelay auth token whenever it needs the key, so the key stays in the deeprelay credentials file. The binary is referenced by its absolute path. If you move or reinstall deeprelay, run setup again.

Where it writes

  • User scope (default): ~/.claude/settings.json, or $CLAUDE_CONFIG_DIR/settings.json if you set that variable. Every project on this machine uses deeprelay.
  • Project scope (--scope project): .claude/settings.local.json in the current directory. Only this project uses deeprelay. Setup warns if git does not ignore that file.

Backups and undo

Before each write, setup copies the current settings file into the backups/claude-code folder of the deeprelay config directory (~/.config/deeprelay, or $XDG_CONFIG_HOME/deeprelay). Backups are timestamped and never overwritten. Setup also records which keys it added, changed or removed, and their earlier values. Your deeprelay key is never in that record; a credential setup removed from the settings file is, so that --remove can restore it.

deeprelay setup claude-code --remove replays that record. Add --scope project to undo a project setup.

  • Keys setup added are deleted.
  • Keys setup changed get their earlier values back.
  • Keys setup removed (a conflicting credential) are put back.
  • Keys you edited yourself after setup are left alone, with a warning.

Warnings

These always print, even with --yes:

  • the settings file already points ANTHROPIC_BASE_URL somewhere else, so setup would redirect all Claude Code usage to deeprelay;
  • the settings file holds a credential or cloud-provider setting that outranks apiKeyHelper, which setup removes (see above);
  • with --scope project, your user settings file holds such a setting; setup does not edit that file, and Claude Code will not reach deeprelay until you remove it there;
  • ANTHROPIC_* variables are set in your shell, which override the settings file (only the names are shown);
  • Claude Code is logged in to a Claude subscription and you chose user scope, so every project routes to deeprelay instead. Consider --scope project.

deeprelay setup claude-code --print writes nothing. It prints export lines for the same variables, with ANTHROPIC_AUTH_TOKEN in place of apiKeyHelper. That output contains your key in plaintext. Don't paste it into a file you commit or a log you share. In CI, prefer a secret store that sets ANTHROPIC_AUTH_TOKEN for the job.

Models

Setup's defaults are models that completed real Claude Code coding sessions (reading files, editing, running commands, thinking, and a follow-up turn with --continue):

SlotDefaultWhy
Main, Opus, Sonnet, Fable (--model, --opus-model, --sonnet-model, --fable-model)deeprelay/deepseek-v4-proThe strongest tested model with thinking and reliable tool calls
Haiku and small-fast (--haiku-model)deeprelay/deepseek-v4.1-flashCovered by the flat plan (a subscriber's background tasks stay inside it), the same family as the main model, and it can see images; used for background tasks such as titles and summaries

deeprelay/glm-5.3 also completed a full session and works as a main model:

deeprelay setup claude-code --model deeprelay/glm-5.3

deeprelay/gpt-oss-120b completed a full session as well. It is not covered by the flat plan, but it is a cheaper pay-as-you-go choice for the background slot: --haiku-model deeprelay/gpt-oss-120b.

Any other model works if it supports tool calling. Setup checks this and refuses a model without it. To see what is available right now:

deeprelay models list

The catalog is live, so a model can come and go. If a model stops being served, run setup again with another one.

Why claude-* model names are rejected

A request for a Claude model ID (for example claude-sonnet-…) gets a 400 instead of being quietly mapped to some other model. deeprelay doesn't alias Claude names, so you always know which model answered, and you are billed for the model you chose. If you see this error, a model slot is not set: run deeprelay setup claude-code, or set the variables below.

Manual setup

If you'd rather not use deeprelay setup, set the same variables yourself, in your shell or in the env block of a Claude Code settings file:

export ANTHROPIC_BASE_URL=https://api.deeprelay.ai
export ANTHROPIC_AUTH_TOKEN=deeprelay_live_...
export ANTHROPIC_MODEL=deeprelay/deepseek-v4-pro
export ANTHROPIC_DEFAULT_OPUS_MODEL=deeprelay/deepseek-v4-pro
export ANTHROPIC_DEFAULT_SONNET_MODEL=deeprelay/deepseek-v4-pro
export ANTHROPIC_DEFAULT_FABLE_MODEL=deeprelay/deepseek-v4-pro
export ANTHROPIC_DEFAULT_HAIKU_MODEL=deeprelay/deepseek-v4.1-flash
export ANTHROPIC_SMALL_FAST_MODEL=deeprelay/deepseek-v4.1-flash
  • ANTHROPIC_BASE_URL has no /v1. Claude Code appends /v1/messages itself, so https://api.deeprelay.ai/v1 would end up at /v1/v1/messages.
  • Set every model variable. A slot you leave unset falls back to a Claude model name, and that request is rejected.
  • Instead of ANTHROPIC_AUTH_TOKEN, you can put "apiKeyHelper": "/absolute/path/to/deeprelay auth token" in the settings file, so the key never sits in your shell or a file.
  • deeprelay accepts the key as x-api-key or as Authorization: Bearer, so ANTHROPIC_API_KEY works too.

What works and what doesn't

Works:

  • Tools. File reads and edits, shell commands, subagents and MCP servers you configure in Claude Code all run through normal tool calls.
  • Streaming, including keep-alive pings during long responses.
  • Thinking. Models that reason stream it as thinking blocks. deeprelay signs them, and handles them correctly when Claude Code sends them back on later turns.
  • Prompt caching. Caching happens automatically on models that support it. cache_control markers are accepted, and cached input shows up as cache_read_input_tokens in usage.

Doesn't work:

  • WebSearch. Claude Code's built-in WebSearch tool relies on an Anthropic server-side tool, which deeprelay does not run. The request returns a 400 that says so, Claude Code shows the error to the model, and the session carries on. For web search, add a search MCP server to Claude Code instead. The same applies to other server tools, such as code execution. WebFetch runs on your machine as an ordinary client tool, so it is not affected.
  • PDFs and documents. Document blocks are rejected with a 400.
  • Images, except on vision models. Today only deeprelay/deepseek-v4.1-flash can see images. On that model, images you paste and images Claude Code's Read tool opens (a screenshot or a PNG in your repo) reach the model. On other models, images never break the session:
    • an image Read opens, or an image from an earlier turn, is replaced by a note telling the model that an image was there and it cannot see it, and the session carries on;
    • only an image you paste into your newest message returns a 400, which names the models that can see images. Remove the image or switch models (/model).

When deeprelay replaces or moves an image, the response carries the header x-deeprelay-ignored-params with image:omitted (replaced by a note) or image:moved (on a vision model, an image from a tool result is passed to the model right after the tool result, because most models accept tool results as text only). Forwarded images must be JPEG, PNG, GIF or WebP, at most 5 MiB each and 20 per request.

  • Exact token counts. POST /v1/messages/count_tokens returns an estimate, marked with the response header x-deeprelay-token-count: estimated. It tends to count high, which is safe for Claude Code's context budgeting. It is free.

Billing

  • A Claude Code request is billed exactly like a chat completion for the same model: the same per-token prices, and the same plan quota and credit.
  • Rejected requests are not billed. That includes a claude-* model name, a web search call and an unsupported image. Token counting is never billed.
  • Usage and spend show the deeprelay model that answered. Check them with deeprelay usage or in the dashboard.
  • Claude Code's /cost does not show deeprelay prices. It computes cost from Anthropic's price list, which doesn't apply to deeprelay models. For what you actually spent, use deeprelay usage or the usage page in the deeprelay dashboard.

Troubleshooting

Errors on /v1/messages use the Anthropic error format, so Claude Code shows them as readable messages. Each error also carries an X-Deeprelay-Error-Code header with the same code the rest of the API uses (see Errors).

"deeprelay doesn't serve Claude models"

deeprelay doesn't serve Claude models — run `deeprelay setup claude-code`, or set ANTHROPIC_MODEL, ANTHROPIC_DEFAULT_{OPUS,SONNET,HAIKU,FABLE}_MODEL and ANTHROPIC_SMALL_FAST_MODEL to deeprelay model IDs (see deeprelay.ai/docs/claude-code).

Claude Code asked for a Claude model, so at least one model slot is not set (code claude_model_not_served). Run deeprelay setup claude-code again, or set every variable listed under Manual setup, including ANTHROPIC_DEFAULT_FABLE_MODEL and ANTHROPIC_SMALL_FAST_MODEL.

"apiKeyHelper is failing"

Claude Code couldn't get a key from the helper. Run the helper yourself to see why:

deeprelay auth token

Usually you are not logged in (run deeprelay login), or the deeprelay binary moved since setup ran (run setup again).

Settings seem to be ignored

  • Shell variables win. An ANTHROPIC_* variable exported in your shell overrides the settings file. Setup lists any it finds. Unset them, or use a clean shell.
  • Claude Code is still using your subscription. If you are logged in to a Claude subscription and ran setup at project scope, only that project uses deeprelay. Check which settings file setup wrote, and run /status in Claude Code to see the active base URL and model.

"prompt is too long"

The conversation no longer fits the model's context window (code context_length_exceeded). deeprelay returns Anthropic's own prompt is too long: N tokens > M maximum error, so Claude Code handles it as it would with Anthropic: it compacts the conversation, or asks you to run /compact. The token count is deeprelay's estimate, and M is the model's context window from deeprelay models get.

402 and 429

StatusX-Deeprelay-Error-CodeMeaning
401unauthorizedThe key is missing, invalid, revoked or expired
402insufficient_balanceOut of credit. Top up in the dashboard
402daily_limit_reached, spending_limit_reachedA spend limit on your account or key was hit
429subscription_quota_exhaustedYour plan's quota for this window is used up
429rate_limit_exceededToo many requests per second. Claude Code retries
429stream_limit_exceededToo many requests open at once on this key. Claude Code retries

Billing 402s and an exhausted plan quota tell Claude Code not to retry, so it reports them once instead of looping. Check where you stand with deeprelay billing subscription and deeprelay usage.

Still stuck? Email support@deeprelay.ai with your Claude Code version, the model ID and the exact error message.

← All docs