Codex CLI
Run Codex CLI on deeprelay models: one setup command writes the provider and profile for the OpenAI Responses endpoint, with what works, what is rejected, and how to undo it.
deeprelay serves the OpenAI Responses API (POST /v1/responses), so
Codex CLI can run on deeprelay models.
Codex keeps working as usual: it reads and edits files, runs commands, calls
tools and compacts long conversations. Every request goes to a deeprelay model
you picked, and it is billed like any other deeprelay inference.
deeprelay does not serve OpenAI models. Setup points Codex at a deeprelay
model, so a configured session never asks for gpt-*, o* or codex-*.
deeprelay is stateless: it stores no responses and no conversations.
Codex already works this way. It sends store: false and the full
conversation on every turn, so nothing about a normal Codex session depends
on storage.
Quick start
deeprelay login # once: stores your key in the deeprelay credentials file
deeprelay setup codex # points Codex at deeprelay (asks before writing)
Setup ends by printing one line for your shell rc (~/.zshrc, ~/.bashrc).
Add it, open a new shell, and start Codex:
export DEEPRELAY_API_KEY="$('/usr/local/bin/deeprelay' auth token)" # your path is printed by setup
codex # start Codex as usual
Setup checks your key and the model, then shows a diff of every file it will change and the undo command, and asks you to confirm. Only after you confirm does it make one small test call, and only if that call works does it write anything.
To try deeprelay without changing Codex's default, use a profile instead:
deeprelay setup codex --profile-only
codex -p deeprelay # or: codex --profile deeprelay
What setup changes
deeprelay setup codex writes two files in the Codex directory (~/.codex/,
or $CODEX_HOME if you set that variable). It edits them one key at a time,
keeps your comments, formatting and every other table exactly as they are,
and never rewrites a file.
config.toml always gets the deeprelay provider:
| Key | Value |
|---|---|
model_providers.deeprelay.name | "deeprelay" |
model_providers.deeprelay.base_url | "https://api.deeprelay.ai/v1" (Codex appends /responses itself) |
model_providers.deeprelay.env_key | "DEEPRELAY_API_KEY", the name of the variable Codex reads the key from |
model_providers.deeprelay.wire_api | "responses" |
Unless you pass --profile-only, config.toml also gets five top-level keys,
which make deeprelay Codex's default:
| Key | Value |
|---|---|
model_provider | "deeprelay" |
model | the model (--model) |
model_reasoning_effort | --effort, or an effort the model supports |
model_reasoning_summary | "auto" |
model_context_window | 95% of the model's context length in the deeprelay catalog, so Codex compacts the conversation before the model runs out of room |
deeprelay.config.toml holds the deeprelay profile as four top-level
keys:
| Key | Value |
|---|---|
model | the model (--model) |
model_provider | "deeprelay" |
model_reasoning_effort | --effort, or an effort the model supports |
model_reasoning_summary | "auto" |
The profile is its own file because current Codex loads
when you run codex --profile ,
and no longer accepts profiles inside config.toml. The effort and summary
are also written at the top of config.toml on the default switch because
plain codex does not read the profile: without them Codex would send no
reasoning effort and would not ask for reasoning summaries.
Setup writes no output-token limit (current Codex has no config key for one)
and never writes a key into http_headers. Every key setup writes is checked
in CI against Codex's own config schema, pinned to the Codex version deeprelay
is tested with.
Your key is never written to a file. env_key holds only the variable
name DEEPRELAY_API_KEY. The export line runs deeprelay auth token each
time a shell starts, which reads your key from the deeprelay credentials file.
If you rotate the key, the next shell picks up the new one without running
setup again. The setup diff shows env_key as because the key
name contains "key"; the value is only the variable name.
Default switch or profile
- Default: the provider and the five top-level keys in
config.toml, plusdeeprelay.config.toml. Start Codex with plaincodex. All Codex usage then goes to deeprelay, including sessions that would otherwise use your ChatGPT plan. --profile-only: the provider inconfig.tomlanddeeprelay.config.toml, nothing else. Codex's default provider, model and reasoning settings stay as they were. Runcodex --profile deeprelay(orcodex -p deeprelay) when you want deeprelay. Codex readsmodel_context_windowonly from the top ofconfig.toml, not from a profile, so the profile uses Codex's default context window; setmodel_context_windowyourself if you need it.
There is no --scope project. Codex also reads a project-level
.codex/config.toml, but that file cannot define providers or profiles. The
deeprelay profile covers the per-project case: run codex -p deeprelay in
the projects where you want deeprelay.
Backups and undo
Before each write, setup copies each file it is about to change to
backups/ in the Codex directory, so both
config.toml and deeprelay.config.toml are backed up. A later run never
overwrites an earlier backup. For each file, setup also records which keys it
added or changed, and their earlier values:
.deeprelay-setup-codex-manifest.json for config.toml and
.deeprelay-setup-codex-profile-manifest.json for the profile file, both in
the same directory. No key is ever in these records.
deeprelay setup codex --remove replays both records key by key:
- Keys setup added are deleted. The
deeprelayprovider table is removed once it is empty, anddeeprelay.config.tomlis deleted once nothing else is in it (a key you added yourself keeps the file). - Keys setup changed get their earlier values back, including your own top-level
model_reasoning_effortormodel_reasoning_summaryif you had them before setup. - Keys you edited yourself after setup are left alone, with a warning naming each key and its file. The undo record is kept so you can sort them out and run
--removeagain.
--remove shows one diff covering both files and asks for confirmation,
like setup. Running it when there is nothing to undo reports "nothing to
remove" and exits 0. After removing, delete the DEEPRELAY_API_KEY line from
your shell rc.
deeprelay setup codex --refresh updates only the keys setup owns from the
current catalog (base_url, the model, model_reasoning_effort and
model_context_window) in both files. It keeps the profile's model unless you
pass --model, keeps your earlier choice between the default switch and
--profile-only, needs an earlier setup, and backs up the files like setup
does. Re-running setup or --refresh when nothing would change prints
"already configured" and makes no test call.
Upgrading from an earlier setup
Setups made before this release wrote the profile as a [profiles.deeprelay]
table inside config.toml. Current Codex refuses that table for -p, and
codex -p deeprelay exits with:
Error loading config.toml: --profile `deeprelay` cannot be used while <codex-home>/config.toml contains legacy `profile = "deeprelay"` or `[profiles.deeprelay]` config; move those settings into <codex-home>/deeprelay.config.toml and remove the legacy profile
Run deeprelay setup codex again (or deeprelay setup codex --refresh). The
diff shows the [profiles.deeprelay] keys that setup wrote being removed and
deeprelay.config.toml being written; after you confirm, setup backs up
config.toml, removes those keys and the empty table, and writes the profile
file. Keys in that table you changed yourself stay where they are: setup
warns about each one and tells you to move it into deeprelay.config.toml or
delete it. Setup does the same, warning only, for a [profiles.deeprelay]
table or a top-level profile = "deeprelay" it did not write.
Warnings
These always print, even with --yes:
- the default switch routes all Codex usage to deeprelay, including ChatGPT-plan sessions. Setup suggests
--profile-onlyinstead; model_provideralready names another provider, which setup replaces;- with
--profile-only, thedeeprelayprofile uses Codex's default context window. Codex readsmodel_context_windowonly at the top level, not from a profile, so Codex may not compact in time on a model with a small window. Setmodel_context_windowyourself, or use the default switch; DEEPRELAY_API_KEYis not set in this shell, followed by the exact line to add to your shell rc;config.tomlstill has[profiles.deeprelay]keys or a top-levelprofile = "deeprelay"that setup does not own (see Upgrading from an earlier setup).
If the file defines model_providers.deeprelay or profiles.deeprelay as
dotted keys, an inline table or an [[array of tables]], setup stops with an
error and writes nothing. Rewrite that part as a normal [table] section and
run setup again.
--print for shells and CI
deeprelay setup codex --print writes nothing and makes no API call. It
prints the export line with the absolute path of your deeprelay binary:
export DEEPRELAY_API_KEY="$('/usr/local/bin/deeprelay' auth token)"
The line contains a command, never the key. It is safe to put in a dotfile
you commit. In CI, skip it and have your secret store set
DEEPRELAY_API_KEY for the job directly.
The binary is referenced by its absolute path, symlinks kept, so a package
manager upgrade does not break it. If you move or reinstall deeprelay
somewhere else, run --print again and update the line.
IDE-launched Codex
Codex started from an IDE, the Dock, the Start menu or another GUI app may
not read your shell rc, so it will not see DEEPRELAY_API_KEY and fails with
a missing-key error. The fix is to make the variable visible to the app that
launches Codex. In every case the key comes from deeprelay auth token, so
never paste the key itself into an IDE settings file: those are saved to
disk, and often synced.
Simplest, every OS: start the IDE from a shell that already has the variable, so it inherits it.
code . # VS Code
idea . # IntelliJ IDEA (or pycharm, goland, webstorm ...)
macOS. GUI apps get their environment from launchd, not your shell.
Set the variable there, then quit and reopen the IDE:
launchctl setenv DEEPRELAY_API_KEY "$(deeprelay auth token)"
This holds the key in memory for your login session, not in a file. It does not survive a reboot or follow a key rotation: run it again after either. Some IDEs also have their own environment settings; if you use one, point it at the variable rather than typing the key.
Linux. Most desktop sessions read ~/.profile at login. Add the export
line printed by setup there, then log out and back in. On systemd-based
desktops you can instead set it for the running user session:
systemctl --user set-environment DEEPRELAY_API_KEY="$(deeprelay auth token)"
Like launchctl, this lives in memory and needs repeating after a reboot or
a key rotation. Don't put the key itself into an environment.d file.
Windows. Codex reads %USERPROFILE%\.codex\config.toml. The cleanest
option is to set the variable in a PowerShell session and start the IDE from
it:
$env:DEEPRELAY_API_KEY = (deeprelay auth token)
code .
To make it permanent for GUI apps, set a user environment variable from a
shell where deeprelay auth token works, then restart the IDE:
setx DEEPRELAY_API_KEY (deeprelay auth token)
setx stores the value in your user environment, so the key itself is
persisted there in plaintext, and it does not follow a key rotation: run it
again after rotating. Prefer the per-session option if that matters to you.
VS Code. The Codex extension runs Codex as a child of VS Code, so it sees
whatever environment VS Code started with: launch VS Code with code . from
a shell, or use the OS-level options above. Codex in the integrated terminal
runs your normal shell, so your rc line applies there. If you use
terminal.integrated.env.osx, .linux or .windows, reference the variable
("${env:DEEPRELAY_API_KEY}") rather than pasting the key into
settings.json.
JetBrains IDEs. On macOS and Linux, JetBrains IDEs normally load your login shell's environment at startup, so the rc line usually just works after a restart. If it doesn't, launch the IDE from a shell or use the OS-level options above. Avoid typing the key into a run configuration or into Settings > Tools > Terminal > Environment variables: both are saved in project or IDE files.
Models
Setup's default is deeprelay/deepseek-v4-pro, from deeprelay's list of
models tested with Codex. Each completed a real multi-turn Codex session with
tool calls and reasoning, and resumed it:
| Model | Notes |
|---|---|
deeprelay/deepseek-v4-pro | The default. The strongest tested model with reasoning and reliable tool calls. Reasoning is sent back to the model on later turns |
deeprelay/deepseek-v4.1-flash | Faster and cheaper. Reasoning is sent back to the model on later turns. Also tested through repeated compaction in a small context window |
deeprelay/glm-5.3 | A different model family. Reasoning shows in Codex but is not sent back to the model on later turns |
Pick another one at setup time, or per session:
deeprelay setup codex --model deeprelay/deepseek-v4.1-flash --effort high
codex -m deeprelay/deepseek-v4.1-flash # this session only
Any other deeprelay model works if it supports tool calling, because Codex does all its work through tools. Setup checks this and refuses a model without it. To see what is available right now:
deeprelay models list
The catalog is live, so a model can come and go. If a model stops being
served, run deeprelay setup codex --model . If its context length
or supported efforts change, deeprelay setup codex --refresh picks up the
new values.
model_reasoning_effort accepts Codex's values: none, minimal, low,
medium, high, xhigh and max. Setup uses --effort if you pass it.
Otherwise it uses medium when the model supports it or the catalog lists no
efforts, and else the first effort the model lists. Change it per session
with codex -c model_reasoning_effort=high.
Codex prints Model metadata for deeprelay/... not found. Defaulting to
fallback metadata on start. That is expected for any model Codex doesn't
know by name, and it is the configuration deeprelay tests with.
With model_reasoning_summary = "auto" in effect (setup writes it in both
files), Codex shows the model's reasoning text as it streams, before the
answer. On models that keep reasoning across turns, the reasoning Codex sends
back is passed to the model again; on the others it is dropped quietly, and
the turn works the same either way.
Why gpt-*, o* and codex-* model names are rejected
A request for an OpenAI model ID (for example gpt-5) gets a 400 with code
model_not_found instead of being quietly mapped to some other model.
deeprelay doesn't alias model names, so you always know which model answered,
and you are billed for the model you chose. The message names the fix:
model `gpt-5` is not served by deeprelay — run `deeprelay setup codex`, or set `model` in ~/.codex/config.toml (or pass -m) to a deeprelay model ID such as deeprelay/deepseek-v4-pro, deeprelay/deepseek-v4.1-flash, deeprelay/glm-5.3
Codex doesn't reformat API errors. It prints the response body as-is, so in
the terminal you see the whole JSON line, printed twice, and the fix is in
message:
ERROR: {"error":{"message":"model `gpt-5` is not served by deeprelay — run `deeprelay setup codex`, …","type":"invalid_request_error","param":"model","code":"model_not_found"}}
Codex makes one request, does not retry, and exits with status 1. If you see
this, either a -m gpt-… flag is overriding your config, or Codex is not
using the deeprelay provider. Run deeprelay setup codex or fix model.
Manual setup
If you'd rather not use deeprelay setup, add this to
~/.codex/config.toml yourself:
# Top-level keys: leave these five out if you only want the profile
model_provider = "deeprelay"
model = "deeprelay/deepseek-v4-pro"
model_reasoning_effort = "medium"
model_reasoning_summary = "auto"
model_context_window = 121600 # example: 95% of a 128,000-token context_length
[model_providers.deeprelay]
name = "deeprelay"
base_url = "https://api.deeprelay.ai/v1"
env_key = "DEEPRELAY_API_KEY"
wire_api = "responses"
then create ~/.codex/deeprelay.config.toml (or
$CODEX_HOME/deeprelay.config.toml) with the profile as top-level keys:
model = "deeprelay/deepseek-v4-pro"
model_provider = "deeprelay"
model_reasoning_effort = "medium"
model_reasoning_summary = "auto"
and this to your shell rc:
export DEEPRELAY_API_KEY="$(deeprelay auth token)"
base_urlincludes/v1. Codex appends/responsesto it.wire_apimust be"responses".- Work out
model_context_windowfrom the model'scontext_lengthindeeprelay models get; the value above is only an example. - With the top-level keys, plain
codexuses deeprelay. Without them, runcodex -p deeprelay. Plaincodexdoes not read the profile file, which is why the effort and summary appear in both places. - Don't put a
[profiles.deeprelay]table inconfig.toml: current Codex refuses-p deeprelaywhile it exists. - Keep
env_keyas a variable name. Don't put a key intohttp_headersor anywhere else inconfig.toml. deeprelay auth tokenreadsDEEPRELAY_API_KEYif it is already set, and otherwise the credentials file written bydeeprelay login. Ifdeeprelayis not on thePATHyour rc sees, use its absolute path.
What works and what doesn't
Works:
- Streaming and non-streaming. Codex streams; the official OpenAI SDK's
responses.create()works both ways. - Function tools. Codex's shell (
exec_command) and its other function tools run as normal function calls. Tool-call ids (call_id) are minted by deeprelay. Tools grouped in anamespace, such as Codex's sub-agent tools inmulti_agent_v1, are flattened into ordinary function tools for the model, and a call to one comes back to Codex under its namespace, so Codex routes it to the right tool. - Reasoning. On a reasoning model you get a
reasoningitem when the request asks for a summary or for the encrypted blob (Codex asks for both), withsummarytext and, when Codex asks for it withinclude: ["reasoning.encrypted_content"], anencrypted_contentblob that Codex sends back on the next turn. The summary is the model's own reasoning, passed through the same privacy scrub as every response. No summarisation is performed and no extra model call is made, soauto,conciseanddetailedall produce the same output. Reasoning sent back from a different model, or with a damaged blob, is dropped quietly rather than rejected. - Structured outputs through
text.format(text,json_object,json_schema). - Images in
input_imageparts, on models with vision support, with the same rules as the rest of the API. - Client-side compaction. Codex compacts a long conversation by sending an ordinary
/v1/responsesrequest with a summarisation prompt, then carries on with a shorter history. That works unchanged; it is one more request, billed like any other turn. In testing, a session squeezed into an 8,000-token window compacted five times and still finished.
Doesn't work (and what you get instead):
| What | Result |
|---|---|
previous_response_id, conversation, background: true | 400, code stateless. Send the full conversation in input |
store: true (or store omitted) | Accepted, nothing is stored. A store: true is listed as store:ignored in the x-deeprelay-ignored-params response header |
GET and DELETE /v1/responses/{id}, POST /v1/responses/{id}/cancel, GET /v1/responses/{id}/input_items | 404, code stateless: deeprelay keeps no responses to retrieve, cancel or delete |
POST /v1/responses/compact, and the context_management request field | 400, code compaction_unsupported. Compact on the client (Codex already does) |
The web_search tool Codex declares on every request | Dropped, not rejected, so turns keep working; the model just has no search tool. Listed as tools.web_search:ignored in x-deeprelay-ignored-params |
Other hosted tools: web_search_preview, file_search, computer_use, image_generation, code_interpreter, mcp | 400, code unsupported_tool_type |
local_shell, shell, apply_patch and custom tool types | 400, code unsupported_tool_type. Declare the tool as a function instead |
tool_choice that forces a hosted tool | 400 |
input_file parts (and input_audio) | 400, code unsupported_content_part |
Input items other than message, function_call, function_call_output and reasoning | 400, code unsupported_item_type |
Other fields OpenAI defines that have no meaning here, such as
parallel_tool_calls, truncation, service_tier or OpenAI's cache-routing
hint, are accepted and ignored. Codex's per-install and per-session identifiers
are never forwarded to the model. Unknown top-level fields are accepted and
ignored too, so a newer Codex that adds one keeps working.
Billing
- A Codex request is billed exactly like a chat completion for the same model: the same per-token prices, and the same plan quota and credit.
- Reasoning tokens are part of the output and billed as output tokens. They are reported in
usage.output_tokens_details.reasoning_tokens. - Each response reports its cost in
usage.cost. A non-streaming response also carries it in thex-deeprelay-costheader; a streamed one (which is what Codex uses) carries it only in theresponse.completedevent. - Codex's compaction requests are ordinary requests and are billed like any other turn.
- Rejected requests are never billed. That includes a foreign model name, a stateless field, a hosted tool and every
/v1/responses/{id}sub-route. - Usage and spend show the deeprelay model that answered. Check them with
deeprelay usageor in the dashboard. Codex's own token counters are an estimate; for what you actually spent, use deeprelay.
Troubleshooting
Errors on /v1/responses use the OpenAI error format. Each error also
carries an X-Deeprelay-Error-Code header with the same code the rest of the
API uses (see Errors). Codex prints the
error body as-is, so the code and message fields are what to read.
"model … is not served by deeprelay"
Codex asked for a model deeprelay doesn't serve (code model_not_found).
See Why gpt-*, o* and codex-* model names are rejected.
Check for a -m flag, a model override in another profile, or a default
that --profile-only left on a non-deeprelay model (with --profile-only,
start Codex with codex -p deeprelay).
"DEEPRELAY_API_KEY is not set" or the export line doesn't resolve
Codex couldn't find the key. Check the variable in the shell you start Codex from:
test -n "$DEEPRELAY_API_KEY" && echo set || echo missing
deeprelay auth token >/dev/null && echo "helper works"
- Missing in a terminal: the rc line isn't there, or the shell was opened before you added it. Run
deeprelay setup codex --print, add the line, and open a new shell. - The helper fails: you are not logged in (run
deeprelay login), or thedeeprelaybinary moved since setup ran (run--printagain and replace the line). - Missing only in an IDE: see IDE-launched Codex.
- A
401with codeunauthorizedmeans a key reached deeprelay but is invalid, revoked or expired. Rundeeprelay loginagain.
Codex says the config is invalid
wire_apimust be"responses"."chat"is rejected by current Codex.- Remove keys current Codex doesn't know. Its error names the offending key; an output-token limit key copied from an old guide is a common one.
- Setup writes only keys validated against Codex's config schema for the version deeprelay is tested with. A much older or newer Codex may name keys differently; update Codex, or run
deeprelay setup codex --removeand set up again. - If setup itself refuses the file because a
deeprelaytable is written as dotted keys or an inline table, rewrite it as a normal[table]section.
"prompt is too long"
The conversation no longer fits the model's context window (code
context_length_exceeded). Codex compacts on its own when it gets close to
model_context_window, which setup sets to 95% of the model's window. If you
hit this anyway:
- you set deeprelay up as a profile only, so Codex is using its default window. Set
model_context_windowat the top level, or run setup without--profile-only; - the model changed since setup ran. Run
deeprelay setup codex --refresh; - run
/compactin Codex, or start a new session.
402, 429 and 503
| Status | X-Deeprelay-Error-Code | Meaning |
|---|---|---|
401 | unauthorized | The key is missing, invalid, revoked or expired |
402 | insufficient_balance | Out of credit. Top up in the dashboard |
402 | daily_limit_reached, spending_limit_reached | A spend limit on your account or key was hit |
429 | subscription_quota_exhausted | Your plan's quota for this window is used up |
429 | rate_limit_exceeded | Too many requests per second. Wait a moment and retry |
429 | stream_limit_exceeded | Too many requests open at once on this key. Wait for one to finish |
503 | upstream_rate_limited, upstream_unavailable | The model's serving capacity is busy or degraded. Transient: wait for Retry-After and retry. A 429 on this endpoint is always about your own key or plan, never about the model's capacity |
Check where you stand with deeprelay billing subscription and
deeprelay usage.
Still stuck? Email support@deeprelay.ai with your Codex version
(codex --version), the model ID and the exact error message.
Related
- Serverless inference API: endpoints, errors and pricing
deeprelay setup codex: every flagdeeprelay auth token: the key helper the export line runs- Claude Code: the same setup for Claude Code
- Coding agents: Cline, Kilo Code, Roo Code, OpenCode, Aider, Continue, Zed and Cursor