deeprelayDocs
CH·GGuides

Codex CLI

Run Codex CLI on deeprelay models: one setup command writes the provider and profile for the OpenAI Responses endpoint, with what works, what is rejected, and how to undo it.

deeprelay serves the OpenAI Responses API (POST /v1/responses), so Codex CLI can run on deeprelay models. Codex keeps working as usual: it reads and edits files, runs commands, calls tools and compacts long conversations. Every request goes to a deeprelay model you picked, and it is billed like any other deeprelay inference.

deeprelay does not serve OpenAI models. Setup points Codex at a deeprelay model, so a configured session never asks for gpt-*, o* or codex-*.

deeprelay is stateless: it stores no responses and no conversations. Codex already works this way. It sends store: false and the full conversation on every turn, so nothing about a normal Codex session depends on storage.

Quick start

deeprelay login                 # once: stores your key in the deeprelay credentials file
deeprelay setup codex           # points Codex at deeprelay (asks before writing)

Setup ends by printing one line for your shell rc (~/.zshrc, ~/.bashrc). Add it, open a new shell, and start Codex:

export DEEPRELAY_API_KEY="$('/usr/local/bin/deeprelay' auth token)"   # your path is printed by setup
codex                           # start Codex as usual

Setup checks your key and the model, then shows a diff of every file it will change and the undo command, and asks you to confirm. Only after you confirm does it make one small test call, and only if that call works does it write anything.

To try deeprelay without changing Codex's default, use a profile instead:

deeprelay setup codex --profile-only
codex -p deeprelay              # or: codex --profile deeprelay

What setup changes

deeprelay setup codex writes two files in the Codex directory (~/.codex/, or $CODEX_HOME if you set that variable). It edits them one key at a time, keeps your comments, formatting and every other table exactly as they are, and never rewrites a file.

config.toml always gets the deeprelay provider:

KeyValue
model_providers.deeprelay.name"deeprelay"
model_providers.deeprelay.base_url"https://api.deeprelay.ai/v1" (Codex appends /responses itself)
model_providers.deeprelay.env_key"DEEPRELAY_API_KEY", the name of the variable Codex reads the key from
model_providers.deeprelay.wire_api"responses"

Unless you pass --profile-only, config.toml also gets five top-level keys, which make deeprelay Codex's default:

KeyValue
model_provider"deeprelay"
modelthe model (--model)
model_reasoning_effort--effort, or an effort the model supports
model_reasoning_summary"auto"
model_context_window95% of the model's context length in the deeprelay catalog, so Codex compacts the conversation before the model runs out of room

deeprelay.config.toml holds the deeprelay profile as four top-level keys:

KeyValue
modelthe model (--model)
model_provider"deeprelay"
model_reasoning_effort--effort, or an effort the model supports
model_reasoning_summary"auto"

The profile is its own file because current Codex loads /.config.toml when you run codex --profile , and no longer accepts profiles inside config.toml. The effort and summary are also written at the top of config.toml on the default switch because plain codex does not read the profile: without them Codex would send no reasoning effort and would not ask for reasoning summaries.

Setup writes no output-token limit (current Codex has no config key for one) and never writes a key into http_headers. Every key setup writes is checked in CI against Codex's own config schema, pinned to the Codex version deeprelay is tested with.

Your key is never written to a file. env_key holds only the variable name DEEPRELAY_API_KEY. The export line runs deeprelay auth token each time a shell starts, which reads your key from the deeprelay credentials file. If you rotate the key, the next shell picks up the new one without running setup again. The setup diff shows env_key as because the key name contains "key"; the value is only the variable name.

Default switch or profile

  • Default: the provider and the five top-level keys in config.toml, plus deeprelay.config.toml. Start Codex with plain codex. All Codex usage then goes to deeprelay, including sessions that would otherwise use your ChatGPT plan.
  • --profile-only: the provider in config.toml and deeprelay.config.toml, nothing else. Codex's default provider, model and reasoning settings stay as they were. Run codex --profile deeprelay (or codex -p deeprelay) when you want deeprelay. Codex reads model_context_window only from the top of config.toml, not from a profile, so the profile uses Codex's default context window; set model_context_window yourself if you need it.

There is no --scope project. Codex also reads a project-level .codex/config.toml, but that file cannot define providers or profiles. The deeprelay profile covers the per-project case: run codex -p deeprelay in the projects where you want deeprelay.

Backups and undo

Before each write, setup copies each file it is about to change to backups/. in the Codex directory, so both config.toml and deeprelay.config.toml are backed up. A later run never overwrites an earlier backup. For each file, setup also records which keys it added or changed, and their earlier values: .deeprelay-setup-codex-manifest.json for config.toml and .deeprelay-setup-codex-profile-manifest.json for the profile file, both in the same directory. No key is ever in these records.

deeprelay setup codex --remove replays both records key by key:

  • Keys setup added are deleted. The deeprelay provider table is removed once it is empty, and deeprelay.config.toml is deleted once nothing else is in it (a key you added yourself keeps the file).
  • Keys setup changed get their earlier values back, including your own top-level model_reasoning_effort or model_reasoning_summary if you had them before setup.
  • Keys you edited yourself after setup are left alone, with a warning naming each key and its file. The undo record is kept so you can sort them out and run --remove again.

--remove shows one diff covering both files and asks for confirmation, like setup. Running it when there is nothing to undo reports "nothing to remove" and exits 0. After removing, delete the DEEPRELAY_API_KEY line from your shell rc.

deeprelay setup codex --refresh updates only the keys setup owns from the current catalog (base_url, the model, model_reasoning_effort and model_context_window) in both files. It keeps the profile's model unless you pass --model, keeps your earlier choice between the default switch and --profile-only, needs an earlier setup, and backs up the files like setup does. Re-running setup or --refresh when nothing would change prints "already configured" and makes no test call.

Upgrading from an earlier setup

Setups made before this release wrote the profile as a [profiles.deeprelay] table inside config.toml. Current Codex refuses that table for -p, and codex -p deeprelay exits with:

Error loading config.toml: --profile `deeprelay` cannot be used while <codex-home>/config.toml contains legacy `profile = "deeprelay"` or `[profiles.deeprelay]` config; move those settings into <codex-home>/deeprelay.config.toml and remove the legacy profile

Run deeprelay setup codex again (or deeprelay setup codex --refresh). The diff shows the [profiles.deeprelay] keys that setup wrote being removed and deeprelay.config.toml being written; after you confirm, setup backs up config.toml, removes those keys and the empty table, and writes the profile file. Keys in that table you changed yourself stay where they are: setup warns about each one and tells you to move it into deeprelay.config.toml or delete it. Setup does the same, warning only, for a [profiles.deeprelay] table or a top-level profile = "deeprelay" it did not write.

Warnings

These always print, even with --yes:

  • the default switch routes all Codex usage to deeprelay, including ChatGPT-plan sessions. Setup suggests --profile-only instead;
  • model_provider already names another provider, which setup replaces;
  • with --profile-only, the deeprelay profile uses Codex's default context window. Codex reads model_context_window only at the top level, not from a profile, so Codex may not compact in time on a model with a small window. Set model_context_window yourself, or use the default switch;
  • DEEPRELAY_API_KEY is not set in this shell, followed by the exact line to add to your shell rc;
  • config.toml still has [profiles.deeprelay] keys or a top-level profile = "deeprelay" that setup does not own (see Upgrading from an earlier setup).

If the file defines model_providers.deeprelay or profiles.deeprelay as dotted keys, an inline table or an [[array of tables]], setup stops with an error and writes nothing. Rewrite that part as a normal [table] section and run setup again.

deeprelay setup codex --print writes nothing and makes no API call. It prints the export line with the absolute path of your deeprelay binary:

export DEEPRELAY_API_KEY="$('/usr/local/bin/deeprelay' auth token)"

The line contains a command, never the key. It is safe to put in a dotfile you commit. In CI, skip it and have your secret store set DEEPRELAY_API_KEY for the job directly.

The binary is referenced by its absolute path, symlinks kept, so a package manager upgrade does not break it. If you move or reinstall deeprelay somewhere else, run --print again and update the line.

IDE-launched Codex

Codex started from an IDE, the Dock, the Start menu or another GUI app may not read your shell rc, so it will not see DEEPRELAY_API_KEY and fails with a missing-key error. The fix is to make the variable visible to the app that launches Codex. In every case the key comes from deeprelay auth token, so never paste the key itself into an IDE settings file: those are saved to disk, and often synced.

Simplest, every OS: start the IDE from a shell that already has the variable, so it inherits it.

code .          # VS Code
idea .          # IntelliJ IDEA (or pycharm, goland, webstorm ...)

macOS. GUI apps get their environment from launchd, not your shell. Set the variable there, then quit and reopen the IDE:

launchctl setenv DEEPRELAY_API_KEY "$(deeprelay auth token)"

This holds the key in memory for your login session, not in a file. It does not survive a reboot or follow a key rotation: run it again after either. Some IDEs also have their own environment settings; if you use one, point it at the variable rather than typing the key.

Linux. Most desktop sessions read ~/.profile at login. Add the export line printed by setup there, then log out and back in. On systemd-based desktops you can instead set it for the running user session:

systemctl --user set-environment DEEPRELAY_API_KEY="$(deeprelay auth token)"

Like launchctl, this lives in memory and needs repeating after a reboot or a key rotation. Don't put the key itself into an environment.d file.

Windows. Codex reads %USERPROFILE%\.codex\config.toml. The cleanest option is to set the variable in a PowerShell session and start the IDE from it:

$env:DEEPRELAY_API_KEY = (deeprelay auth token)
code .

To make it permanent for GUI apps, set a user environment variable from a shell where deeprelay auth token works, then restart the IDE:

setx DEEPRELAY_API_KEY (deeprelay auth token)

setx stores the value in your user environment, so the key itself is persisted there in plaintext, and it does not follow a key rotation: run it again after rotating. Prefer the per-session option if that matters to you.

VS Code. The Codex extension runs Codex as a child of VS Code, so it sees whatever environment VS Code started with: launch VS Code with code . from a shell, or use the OS-level options above. Codex in the integrated terminal runs your normal shell, so your rc line applies there. If you use terminal.integrated.env.osx, .linux or .windows, reference the variable ("${env:DEEPRELAY_API_KEY}") rather than pasting the key into settings.json.

JetBrains IDEs. On macOS and Linux, JetBrains IDEs normally load your login shell's environment at startup, so the rc line usually just works after a restart. If it doesn't, launch the IDE from a shell or use the OS-level options above. Avoid typing the key into a run configuration or into Settings > Tools > Terminal > Environment variables: both are saved in project or IDE files.

Models

Setup's default is deeprelay/deepseek-v4-pro, from deeprelay's list of models tested with Codex. Each completed a real multi-turn Codex session with tool calls and reasoning, and resumed it:

ModelNotes
deeprelay/deepseek-v4-proThe default. The strongest tested model with reasoning and reliable tool calls. Reasoning is sent back to the model on later turns
deeprelay/deepseek-v4.1-flashFaster and cheaper. Reasoning is sent back to the model on later turns. Also tested through repeated compaction in a small context window
deeprelay/glm-5.3A different model family. Reasoning shows in Codex but is not sent back to the model on later turns

Pick another one at setup time, or per session:

deeprelay setup codex --model deeprelay/deepseek-v4.1-flash --effort high
codex -m deeprelay/deepseek-v4.1-flash          # this session only

Any other deeprelay model works if it supports tool calling, because Codex does all its work through tools. Setup checks this and refuses a model without it. To see what is available right now:

deeprelay models list

The catalog is live, so a model can come and go. If a model stops being served, run deeprelay setup codex --model . If its context length or supported efforts change, deeprelay setup codex --refresh picks up the new values.

model_reasoning_effort accepts Codex's values: none, minimal, low, medium, high, xhigh and max. Setup uses --effort if you pass it. Otherwise it uses medium when the model supports it or the catalog lists no efforts, and else the first effort the model lists. Change it per session with codex -c model_reasoning_effort=high.

Codex prints Model metadata for deeprelay/... not found. Defaulting to fallback metadata on start. That is expected for any model Codex doesn't know by name, and it is the configuration deeprelay tests with.

With model_reasoning_summary = "auto" in effect (setup writes it in both files), Codex shows the model's reasoning text as it streams, before the answer. On models that keep reasoning across turns, the reasoning Codex sends back is passed to the model again; on the others it is dropped quietly, and the turn works the same either way.

Why gpt-*, o* and codex-* model names are rejected

A request for an OpenAI model ID (for example gpt-5) gets a 400 with code model_not_found instead of being quietly mapped to some other model. deeprelay doesn't alias model names, so you always know which model answered, and you are billed for the model you chose. The message names the fix:

model `gpt-5` is not served by deeprelay — run `deeprelay setup codex`, or set `model` in ~/.codex/config.toml (or pass -m) to a deeprelay model ID such as deeprelay/deepseek-v4-pro, deeprelay/deepseek-v4.1-flash, deeprelay/glm-5.3

Codex doesn't reformat API errors. It prints the response body as-is, so in the terminal you see the whole JSON line, printed twice, and the fix is in message:

ERROR: {"error":{"message":"model `gpt-5` is not served by deeprelay — run `deeprelay setup codex`, …","type":"invalid_request_error","param":"model","code":"model_not_found"}}

Codex makes one request, does not retry, and exits with status 1. If you see this, either a -m gpt-… flag is overriding your config, or Codex is not using the deeprelay provider. Run deeprelay setup codex or fix model.

Manual setup

If you'd rather not use deeprelay setup, add this to ~/.codex/config.toml yourself:

# Top-level keys: leave these five out if you only want the profile
model_provider = "deeprelay"
model = "deeprelay/deepseek-v4-pro"
model_reasoning_effort = "medium"
model_reasoning_summary = "auto"
model_context_window = 121600   # example: 95% of a 128,000-token context_length

[model_providers.deeprelay]
name = "deeprelay"
base_url = "https://api.deeprelay.ai/v1"
env_key = "DEEPRELAY_API_KEY"
wire_api = "responses"

then create ~/.codex/deeprelay.config.toml (or $CODEX_HOME/deeprelay.config.toml) with the profile as top-level keys:

model = "deeprelay/deepseek-v4-pro"
model_provider = "deeprelay"
model_reasoning_effort = "medium"
model_reasoning_summary = "auto"

and this to your shell rc:

export DEEPRELAY_API_KEY="$(deeprelay auth token)"
  • base_url includes /v1. Codex appends /responses to it.
  • wire_api must be "responses".
  • Work out model_context_window from the model's context_length in deeprelay models get ; the value above is only an example.
  • With the top-level keys, plain codex uses deeprelay. Without them, run codex -p deeprelay. Plain codex does not read the profile file, which is why the effort and summary appear in both places.
  • Don't put a [profiles.deeprelay] table in config.toml: current Codex refuses -p deeprelay while it exists.
  • Keep env_key as a variable name. Don't put a key into http_headers or anywhere else in config.toml.
  • deeprelay auth token reads DEEPRELAY_API_KEY if it is already set, and otherwise the credentials file written by deeprelay login. If deeprelay is not on the PATH your rc sees, use its absolute path.

What works and what doesn't

Works:

  • Streaming and non-streaming. Codex streams; the official OpenAI SDK's responses.create() works both ways.
  • Function tools. Codex's shell (exec_command) and its other function tools run as normal function calls. Tool-call ids (call_id) are minted by deeprelay. Tools grouped in a namespace, such as Codex's sub-agent tools in multi_agent_v1, are flattened into ordinary function tools for the model, and a call to one comes back to Codex under its namespace, so Codex routes it to the right tool.
  • Reasoning. On a reasoning model you get a reasoning item when the request asks for a summary or for the encrypted blob (Codex asks for both), with summary text and, when Codex asks for it with include: ["reasoning.encrypted_content"], an encrypted_content blob that Codex sends back on the next turn. The summary is the model's own reasoning, passed through the same privacy scrub as every response. No summarisation is performed and no extra model call is made, so auto, concise and detailed all produce the same output. Reasoning sent back from a different model, or with a damaged blob, is dropped quietly rather than rejected.
  • Structured outputs through text.format (text, json_object, json_schema).
  • Images in input_image parts, on models with vision support, with the same rules as the rest of the API.
  • Client-side compaction. Codex compacts a long conversation by sending an ordinary /v1/responses request with a summarisation prompt, then carries on with a shorter history. That works unchanged; it is one more request, billed like any other turn. In testing, a session squeezed into an 8,000-token window compacted five times and still finished.

Doesn't work (and what you get instead):

WhatResult
previous_response_id, conversation, background: true400, code stateless. Send the full conversation in input
store: true (or store omitted)Accepted, nothing is stored. A store: true is listed as store:ignored in the x-deeprelay-ignored-params response header
GET and DELETE /v1/responses/{id}, POST /v1/responses/{id}/cancel, GET /v1/responses/{id}/input_items404, code stateless: deeprelay keeps no responses to retrieve, cancel or delete
POST /v1/responses/compact, and the context_management request field400, code compaction_unsupported. Compact on the client (Codex already does)
The web_search tool Codex declares on every requestDropped, not rejected, so turns keep working; the model just has no search tool. Listed as tools.web_search:ignored in x-deeprelay-ignored-params
Other hosted tools: web_search_preview, file_search, computer_use, image_generation, code_interpreter, mcp400, code unsupported_tool_type
local_shell, shell, apply_patch and custom tool types400, code unsupported_tool_type. Declare the tool as a function instead
tool_choice that forces a hosted tool400
input_file parts (and input_audio)400, code unsupported_content_part
Input items other than message, function_call, function_call_output and reasoning400, code unsupported_item_type

Other fields OpenAI defines that have no meaning here, such as parallel_tool_calls, truncation, service_tier or OpenAI's cache-routing hint, are accepted and ignored. Codex's per-install and per-session identifiers are never forwarded to the model. Unknown top-level fields are accepted and ignored too, so a newer Codex that adds one keeps working.

Billing

  • A Codex request is billed exactly like a chat completion for the same model: the same per-token prices, and the same plan quota and credit.
  • Reasoning tokens are part of the output and billed as output tokens. They are reported in usage.output_tokens_details.reasoning_tokens.
  • Each response reports its cost in usage.cost. A non-streaming response also carries it in the x-deeprelay-cost header; a streamed one (which is what Codex uses) carries it only in the response.completed event.
  • Codex's compaction requests are ordinary requests and are billed like any other turn.
  • Rejected requests are never billed. That includes a foreign model name, a stateless field, a hosted tool and every /v1/responses/{id} sub-route.
  • Usage and spend show the deeprelay model that answered. Check them with deeprelay usage or in the dashboard. Codex's own token counters are an estimate; for what you actually spent, use deeprelay.

Troubleshooting

Errors on /v1/responses use the OpenAI error format. Each error also carries an X-Deeprelay-Error-Code header with the same code the rest of the API uses (see Errors). Codex prints the error body as-is, so the code and message fields are what to read.

"model … is not served by deeprelay"

Codex asked for a model deeprelay doesn't serve (code model_not_found). See Why gpt-*, o* and codex-* model names are rejected. Check for a -m flag, a model override in another profile, or a default that --profile-only left on a non-deeprelay model (with --profile-only, start Codex with codex -p deeprelay).

"DEEPRELAY_API_KEY is not set" or the export line doesn't resolve

Codex couldn't find the key. Check the variable in the shell you start Codex from:

test -n "$DEEPRELAY_API_KEY" && echo set || echo missing
deeprelay auth token >/dev/null && echo "helper works"
  • Missing in a terminal: the rc line isn't there, or the shell was opened before you added it. Run deeprelay setup codex --print, add the line, and open a new shell.
  • The helper fails: you are not logged in (run deeprelay login), or the deeprelay binary moved since setup ran (run --print again and replace the line).
  • Missing only in an IDE: see IDE-launched Codex.
  • A 401 with code unauthorized means a key reached deeprelay but is invalid, revoked or expired. Run deeprelay login again.

Codex says the config is invalid

  • wire_api must be "responses". "chat" is rejected by current Codex.
  • Remove keys current Codex doesn't know. Its error names the offending key; an output-token limit key copied from an old guide is a common one.
  • Setup writes only keys validated against Codex's config schema for the version deeprelay is tested with. A much older or newer Codex may name keys differently; update Codex, or run deeprelay setup codex --remove and set up again.
  • If setup itself refuses the file because a deeprelay table is written as dotted keys or an inline table, rewrite it as a normal [table] section.

"prompt is too long"

The conversation no longer fits the model's context window (code context_length_exceeded). Codex compacts on its own when it gets close to model_context_window, which setup sets to 95% of the model's window. If you hit this anyway:

  • you set deeprelay up as a profile only, so Codex is using its default window. Set model_context_window at the top level, or run setup without --profile-only;
  • the model changed since setup ran. Run deeprelay setup codex --refresh;
  • run /compact in Codex, or start a new session.

402, 429 and 503

StatusX-Deeprelay-Error-CodeMeaning
401unauthorizedThe key is missing, invalid, revoked or expired
402insufficient_balanceOut of credit. Top up in the dashboard
402daily_limit_reached, spending_limit_reachedA spend limit on your account or key was hit
429subscription_quota_exhaustedYour plan's quota for this window is used up
429rate_limit_exceededToo many requests per second. Wait a moment and retry
429stream_limit_exceededToo many requests open at once on this key. Wait for one to finish
503upstream_rate_limited, upstream_unavailableThe model's serving capacity is busy or degraded. Transient: wait for Retry-After and retry. A 429 on this endpoint is always about your own key or plan, never about the model's capacity

Check where you stand with deeprelay billing subscription and deeprelay usage.

Still stuck? Email support@deeprelay.ai with your Codex version (codex --version), the model ID and the exact error message.

← All docs