Serverless Inference API
The OpenAI-compatible /v1 API in full: chat, embeddings, image and video, the model catalog, tool calling, the Anthropic Messages endpoint, plan coverage, and errors.
deeprelay's serverless inference is an OpenAI-compatible HTTP API: point any
OpenAI SDK (or curl, or the deeprelay CLI) at it and call chat completions,
embeddings, image generation, video generation, and the model catalog. You pay
per token (or per image / per video-second) — no instances to manage.
Base URL & authentication
| Base URL | https://api.deeprelay.ai/v1 |
| Auth | Authorization: Bearer deeprelay_live_… |
Get a key with deeprelay login (it stores a deeprelay_live_… key in
~/.config/deeprelay/credentials.json) or mint one in the dashboard. Scopes:
serverless:read lists models, serverless:write runs
chat/embeddings/image/video,
billing:read reads usage; a full_access key covers all three.
Every key you mint belongs to this production host. api.demo.deeprelay.ai,
which appears in some older examples, is our internal staging environment
with its own key store, so a production key gets a 401 there. /v1 is the
API version and the only one there is; the OpenAPI document's 1.0.0 is the
revision of that /v1 contract.
Trial keys
If someone at deeprelay sent you a trial key, it is a normal deeprelay_live_…
API key with trial credit already loaded — no account or sign-in needed. Use it
exactly like any other key against the base URL above:
export DEEPRELAY_API_KEY=deeprelay_live_… # the key you were sent
curl https://api.deeprelay.ai/v1/chat/completions \
-H "Authorization: Bearer $DEEPRELAY_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"deeprelay/deepseek-v4-pro","messages":[{"role":"user","content":"Hello!"}]}'
Or with an OpenAI SDK, set base_url="https://api.deeprelay.ai/v1" and api_key
to the trial key. GET /v1/models lists every model and its per-token price;
the DeepSeek, GLM and Kimi families are the ones the trial is meant for.
What a trial key can and cannot do:
- Scope: serverless inference (
serverless:write) plus reading its own balance (billing:read). It cannot launch GPU instances or mint other keys. - Credit is the cap.
GET /v1/billing/balanceshows what is left. When the balance reaches zero, requests return402 insufficient_balance. - It expires on the date in the message you received (30 days by default). After that, inference calls (
/v1/chat/completions,/v1/modelsand the other OpenAI-compatible routes) return401with the OpenAI-style error codeunauthorized;GET /v1/billing/balancereturns the RFC 7807 problem codeinvalid_api_key. - Rate limit is 10 requests/second, enough for trying models by hand.
- To keep going after the trial, sign up at deeprelay.ai and mint your own key — trial credit does not transfer.
Plan coverage
Models the flat subscription tier covers carry "plan_covered": true in
GET /v1/models (and a PLAN/Plan marker in deeprelay models list /
deeprelay models get) — a subscriber's requests to those models draw on the plan's
monthly allowance first; everything else bills pay-as-you-go.
GET /v1/billing/subscription (CLI: deeprelay billing subscription) reports
how much of that allowance is left.
Being subscribed does not make every model free, and the response to a request
for an uncovered model does not mention that it was charged. So ask first:
GET /v1/inference/preflight?model= returns ok / warn / block for
that exact model, computed from the same gates that will judge the request —
plan coverage, entitlement, remaining quota, credit and spending caps. The
deeprelay CLI runs it automatically before chat, embeddings, image and
video create, warning when a call will bill pay-as-you-go and refusing when
it would be denied; deeprelay preflight asks it directly, and
--no-preflight turns the automatic check off.
Tiers — serverless (default) and economy
A model can have two tiers:
- serverless (default) — low latency, no cold start. Use the bare model id, e.g.
deeprelay/deepseek-v4-flash. - economy — cheaper, compute-priced, may cold-start ~30–60s. Opt in with the
:economysuffix, e.g.deeprelay/.:economy
Economy exists only where GET /v1/models lists the suffixed id as its own
entry. Check the live catalog before relying on it: a :economy id the
catalog does not list returns model_not_found, and there are times when no
model has an economy tier at all.
deeprelay models list shows both tiers as separate rows with their own pricing.
The serverless catalog is live
The serverless catalog is built from the models our upstream partners are
actually serving warm right now — not a hand-maintained list. It is
refreshed roughly hourly and probe-verified, so GET /v1/models always reflects
what you can call this minute. Each entry carries name, author (the model's
creator — e.g. Meta, Qwen, Black Forest Labs), category (mirrors
modality: chat, embeddings, image, or video), and pricing (the cheaper
option when a model is served by more than one partner). Chat
models are priced per 1M tokens (input_per_1m_tokens_cents /
output_per_1m_tokens_cents), plus an optional discounted rate for the
cache-hit part of the prompt (cached_input_per_1m_tokens_microcents — see
Prompt caching below). Embedding models are priced per 1M input
tokens only (input_per_1m_tokens_cents); they emit no output tokens, so
output_per_1m_tokens_cents is absent. Image models are priced per output
megapixel when the partner meters that way (per_mpxl_microcents, in
micro-cents — e.g. 270000 = $0.0027/Mpxl), otherwise with a representative
per-image "from" price (per_image_microcents). Video models are billed per
second of output video (per_video_second_microcents); the per-clip
per_video_microcents value is the representative "from" price for the
model's default clip. The set of authors we surface grows over time.
Because the catalog is live, a model can occasionally drop out between refreshes.
If you call one that just became unavailable you'll get a clear
model_not_found error asking you to try another — see Errors below.
Using the deeprelay CLI
# Discover models
deeprelay models list --modality chat
deeprelay models get deeprelay/deepseek-v4-flash
# Chat (streams tokens)
deeprelay chat deeprelay/deepseek-v4-flash "Explain WireGuard in one sentence"
# Embeddings (one vector per input; -o json for the full vectors)
deeprelay embeddings deeprelay/qwen3-embedding-8b "first sentence" "second sentence"
# Image generation (writes out.png)
deeprelay image "a watercolor fox in a misty forest" --model deeprelay/flux.2-dev
# Video generation (async: submit, poll, download)
deeprelay video create --model deeprelay/seedance-1.0-lite --prompt "ocean waves at sunset" --wait 10m
deeprelay video download <video-id> --out waves.mp4
# Usage / spend
deeprelay usage --modality chat --bucket month
# On the flat tier: status plus how much quota is left
deeprelay billing subscription
# Will this model draw on the plan, or on credit?
deeprelay preflight deeprelay/deepseek-v4-flash
# Start or cancel a subscription (each prints a link to open)
deeprelay billing subscribe
deeprelay billing cancel
# Refer a friend: you both get $5 in credit once they spend $5
deeprelay referral
See the per-command reference: chat, embeddings, image, video, models, usage, preflight, billing subscription, billing subscribe, billing cancel, referral (see the refer-a-friend guide).
OpenAI compatibility
The chat request and response shapes, streaming, tool calling and the error
envelope match OpenAI's. Not every OpenAI chat parameter is carried, though.
Each one is forwarded to the model, rejected with a 400 before
anything is billed, or ignored (accepted and dropped):
| Parameter | Handling |
|---|---|
model, messages | Forwarded. content may be a string, null, or an array of text parts ([{"type":"text","text":"…"}], joined with newlines). Roles are system, user, assistant and tool; developer is treated as system. |
temperature, top_p, max_tokens, stop, n, seed, user | Forwarded. |
max_completion_tokens | Forwarded as max_tokens. If you send both, max_tokens wins. |
stream, stream_options | Forwarded. |
tools, tool_choice | Forwarded, on models that list tools in supported_parameters (see Tool calling). finish_reason is tool_calls when the model makes a call. |
response_format, logprobs, top_logprobs, logit_bias, function_call | Rejected: 400 unsupported_parameter (see JSON mode). |
image_url, input_audio, file content parts | Rejected: 400 unsupported_content_part, param: "messages". No model in the catalog takes image, audio or file input. |
reasoning_effort, presence_penalty, frequency_penalty, parallel_tool_calls, anything else | Ignored: accepted without error and not sent to the model. |
Ignored means not applied. reasoning_effort is the one most likely to
matter: a reasoning model thinks at its own default effort whatever you send,
and its reasoning still streams as reasoning deltas. Don't rely on an ignored
parameter for correctness.
Using the OpenAI SDK
Because the API is OpenAI-compatible, the official SDKs work unchanged for
chat (within the table above). Just override base_url and api_key:
from openai import OpenAI
client = OpenAI(
base_url="https://api.deeprelay.ai/v1",
api_key="deeprelay_live_…",
)
# Chat (streaming)
stream = client.chat.completions.create(
model="deeprelay/deepseek-v4-flash",
messages=[{"role": "user", "content": "Explain WireGuard in one sentence"}],
stream=True,
)
for chunk in stream:
print(chunk.choices[0].delta.content or "", end="")
Coding agents
The same base URL and key work in coding agents that accept a custom OpenAI-compatible provider, such as Cline, Kilo Code, Roo Code, OpenCode, Aider, Continue, Zed and Cursor. Coding agents has the config for each and the models to pick for tool calling.
Claude Code uses the Anthropic Messages API instead. Run
deeprelay setup claude-code, and see Claude Code and
Anthropic Messages API below.
Using curl
# List models
curl https://api.deeprelay.ai/v1/models \
-H "Authorization: Bearer deeprelay_live_…"
# Chat completion (non-streaming)
curl https://api.deeprelay.ai/v1/chat/completions \
-H "Authorization: Bearer deeprelay_live_…" \
-H "Content-Type: application/json" \
-d '{"model":"deeprelay/deepseek-v4-flash","messages":[{"role":"user","content":"hi"}]}'
Anthropic Messages API
deeprelay also serves the Anthropic Messages API at POST /v1/messages, so
Claude Code and the Anthropic SDKs can use deeprelay models. For Claude Code,
run deeprelay setup claude-code. Claude Code is the full
guide. The Anthropic SDKs take the base URL without /v1
(https://api.deeprelay.ai).
curl https://api.deeprelay.ai/v1/messages \
-H "x-api-key: deeprelay_live_…" \
-H "anthropic-version: 2023-06-01" \
-H "Content-Type: application/json" \
-d '{"model":"deeprelay/deepseek-v4-pro","max_tokens":256,"messages":[{"role":"user","content":"hi"}]}'
- Auth: send the key as
x-api-keyor asAuthorization: Bearer. - Models: use deeprelay model IDs. A
claude-*model ID is rejected with400(claude_model_not_served) and a message telling you to rundeeprelay setup claude-code. deeprelay doesn't alias Claude names to other models. - Supported: text, tools, streaming (with
pingkeep-alives), thinking and prompt caching.cache_controlmarkers are accepted, and caching happens automatically where the model supports it. - Not supported, each rejected with a
400:- server-side tools such as web search, web fetch and code execution (
server_tool_unsupported). Use an MCP server in the client instead; mcp_serversandcontainer(server_side_feature_unsupported);- document and PDF blocks;
- a user image in the newest user message on a model without vision support (
unsupported_capability; the message names the models that accept images). Today onlydeeprelay/deepseek-v4.1-flashaccepts images on this endpoint.
- server-side tools such as web search, web fetch and code execution (
- Images that are not rejected: on a model without vision, an image in an earlier turn or inside a
tool_resultis replaced by a text note saying an image was omitted. On a vision model, an image inside atool_resultis moved into user content right after the tool result, because most models take tool results as text only. The response headerx-deeprelay-ignored-paramsthen listsimage:omittedand/orimage:moved. - Stop sequences:
stop_sequencesare honoured, but when one ends the response the result saysstop_reason: "end_turn"withstop_sequence: null, because the models behind this endpoint do not report which sequence matched. - Token counting:
POST /v1/messages/count_tokensreturns an estimate, marked with the headerx-deeprelay-token-count: estimated. It is never billed. - Errors use the Anthropic envelope,
{"type":"error","error":{"type":"…","message":"…"},"request_id":"…"}. The deeprelay code rides in theX-Deeprelay-Error-Codeheader, from the same set of codes as the rest of the API (see Errors). - Billing is the same as a chat completion on the same model. Rejected requests are not billed.
Endpoints
| Method | Path | Scope | Purpose |
|---|---|---|---|
GET | /v1/models (?modality=) | serverless:read | List the model catalog |
GET | /v1/models/{id} | serverless:read | One model's record |
POST | /v1/chat/completions | serverless:write | Chat completion (supports stream:true) |
POST | /v1/messages | serverless:write | Anthropic Messages API, for Claude Code (supports stream:true); see Anthropic Messages API |
POST | /v1/messages/count_tokens | serverless:read | Estimated input-token count for a Messages request (local estimate, not billed) |
POST | /v1/embeddings | serverless:write | Text embeddings (OpenAI-compatible, synchronous) |
POST | /v1/images/generations | serverless:write | Synchronous image generation (b64_json) |
POST | /v1/videos | serverless:write | Create an async video job |
GET | /v1/videos (?after= / ?limit=) | serverless:read | List your video jobs |
GET | /v1/videos/{id} | serverless:read | One video job (poll for status) |
GET | /v1/videos/{id}/content | serverless:read | Download the finished mp4 |
POST | /v1/videos/{id}/cancel | serverless:write | Cancel an in-flight video job |
GET | /v1/usage (?modality= / ?model=) | billing:read | Token/cost rollups |
GET | /v1/billing/subscription | billing:read | Flat-tier status, plan + price, quota usage (200 with subscribed:false when not subscribed) |
POST | /v1/billing/subscription/checkout | billing:write (org admin) | Mint a Stripe Checkout URL to subscribe (409 if already subscribed) |
POST | /v1/billing/subscription/portal | billing:write (org admin) | Mint a Stripe Billing Portal URL to cancel, resume or change card |
GET | /v1/inference/preflight (?model=) | serverless:read | Would this model request be served, and at whose expense |
Reading your usage and spend
GET /v1/usage returns your inference spend in time buckets (bucket=hour,
day, week or month; default day) over start/end (RFC 3339; default
the last 30 days). Pass modality and/or model — those filters are what
select serverless inference usage. Without either, the endpoint reads GPU
instance usage, which is always empty today.
# Chat spend per day for the last 30 days
curl "https://api.deeprelay.ai/v1/usage?modality=chat" \
-H "Authorization: Bearer $DEEPRELAY_API_KEY"
# One model, hourly, over an explicit window
curl "https://api.deeprelay.ai/v1/usage?model=<model-id>&bucket=hour&start=2026-09-01T00:00:00Z&end=2026-09-02T00:00:00Z" \
-H "Authorization: Bearer $DEEPRELAY_API_KEY"
modality is one of chat, image, video or embedding; model is an
exact model id from GET /v1/models. Each row covers one (bucket, modality,
model) and carries bucket_start, modality, model, prompt_tokens,
completion_tokens, image_count and cost_cents (plus gpu_seconds, always
0 on inference rows). Rows come newest bucket first, paginated with
limit/cursor like every other list. The CLI equivalent is
deeprelay usage --modality chat (or --model ), and the SDKs expose both
filters on listUsage / list_usage.
Prompt caching (cached input tokens)
When an upstream partner serves part of your prompt from its prompt cache, we pass the split straight through and bill the cached part at a lower rate.
In the response. A chat completion's usage block carries
prompt_tokens_details.cached_tokens whenever the upstream reported a non-zero
cache hit — the same field name and shape the OpenAI SDKs already read, on both
streaming and non-streaming calls:
"usage": {
"prompt_tokens": 4096,
"completion_tokens": 128,
"total_tokens": 4224,
"prompt_tokens_details": { "cached_tokens": 3840 }
}
cached_tokens is a subset of prompt_tokens, never an addition:
prompt_tokens remains the full prompt. When there were no cache hits the
prompt_tokens_details object is omitted entirely.
In the catalog. A model with a cached-input rate exposes it in pricing as
cached_input_per_1m_tokens_microcents (micro-cents per 1M tokens, the exact
figure — cached rates are routinely sub-cent) alongside a rounded
cached_input_per_1m_tokens_cents. Treat the micro-cents field as the
presence signal: it is absent only when the model has no cached rate (no
discount — every prompt token bills at input_per_1m_tokens_cents), while the
rounded cents field is also omitted whenever the exact rate rounds below half
a cent, which for cached rates is the routine case.
How a chat call is billed.
cost = (prompt_tokens - cached_tokens) × input_rate
+ cached_tokens × cached_input_rate
+ completion_tokens × output_rate
rounded up to the next whole cent per request, with all rates per 1M tokens.
cached_tokens is clamped to prompt_tokens, so an over-reporting upstream can
never produce a negative cache-miss count. Prompt caching is a partner-side
optimisation — we don't control which prefixes get cached, and a call that hits
no cache simply bills exactly as it did before.
Where the cached rate comes from. A model's cached-input rate is taken from
the serving partner's published cache-hit price (the same listing its input and
output rates come from) and carries the same retail markup, or from the pricing
policy set for the model when one is enrolled. It is visible everywhere the
model's price is: GET /v1/models, deeprelay models list / deeprelay models get, the
public model pricing pages and the console's model
catalog. A model that shows no cached rate has none — its partner publishes no
cache-hit price — and bills every prompt token at the input rate.
Peak and off-peak pricing
A few models are priced by time of day: their upstream charges more inside
daily peak hours. For those models pricing carries the whole picture:
"pricing": {
"currency": "usd",
"input_per_1m_tokens_cents": 30,
"output_per_1m_tokens_cents": 89,
"cached_input_per_1m_tokens_microcents": 114000,
"period": "peak",
"peak_windows_utc": [
{"start": "01:00", "end": "04:00"},
{"start": "06:00", "end": "10:00"}
],
"peak": {"input_per_1m_tokens_microcents": 30180000, "cached_input_per_1m_tokens_microcents": 114000, "output_per_1m_tokens_microcents": 88620000, "...": "..."},
"off_peak": {"input_per_1m_tokens_microcents": 15090000, "cached_input_per_1m_tokens_microcents": 57000, "output_per_1m_tokens_microcents": 44310000, "...": "..."}
}
- The flat fields (
input_per_1m_tokens_cents,output_per_1m_tokens_cents,cached_input_per_1m_tokens_microcents) are the rates in effect at the moment you read the catalog, andperiodsays which side of the schedule that is (peakoroff_peak). A client that reads only the flat fields always sees the price a request sent right now is billed at. peak_windows_utcis the daily schedule in UTC — start inclusive, end exclusive,"24:00"allowed as an end, and a range whose end sorts before its start wraps midnight.peakandoff_peakcarry both full rate triples so you can price a call you will make later.- A request is billed at the rate in effect when it arrives (its UTC arrival minute is checked against the windows). A long streamed response that straddles a boundary is priced entirely at the arrival rate.
- Models without time-of-day pricing omit all four fields; their flat rates apply around the clock.
deeprelay models get shows the period, the windows and both triples for such
a model.
Balance checks and usage rollups. The pre-flight balance check assumes
zero cache hits, so it stays conservative: a cached call can only cost less
than the amount reserved. GET /v1/usage and the Stripe input-token meter keep
counting the full prompt_tokens — the cache discount shows up in the cost,
not in the token counts.
Image generation
POST /v1/images/generations is synchronous and OpenAI-compatible. v1 returns
images as base64 only — data[].b64_json carries the bytes, and
response_format:"url" is rejected with an invalid_request_error. The response
adds a usage.image_count field (the metered unit).
curl https://api.deeprelay.ai/v1/images/generations \
-H "Authorization: Bearer deeprelay_live_…" \
-H "Content-Type: application/json" \
-d '{"model":"deeprelay/flux.2-dev","prompt":"a watercolor fox","n":1,"size":"1024x1024"}'
{ "created": 1733800000, "data": [{ "b64_json": "iVBORw0KGgo…" }], "usage": { "image_count": 1 } }
The deeprelay image CLI decodes the base64 and writes the file(s) for you:
deeprelay image "a watercolor fox in a misty forest" --model deeprelay/flux.2-dev --out fox.png
Video generation
Video generation is asynchronous. POST /v1/videos returns a job in the
queued state immediately; poll GET /v1/videos/{id} until the status reaches a
terminal value, then download the mp4 from GET /v1/videos/{id}/content.
Lifecycle. A job moves queued → in_progress → completed (or failed /
cancelled / expired). While in_progress the job carries a progress
value (0–100); some models report no granular progress and stay at 0 until
they complete.
Parameters. Each model declares the request parameters it accepts in its
catalog entry (parameters on GET /v1/models) — typically seconds (clip
length; bounds and default come from the descriptor). Parameters a model does
not declare are rejected with an invalid_request_error rather than silently
ignored: live video models generate at their default output resolution, so
most do not accept size. Omitting seconds uses the model's default clip
length.
Failures. A failed job carries a canonical code in error
(invalid_parameters, content_policy, upstream_error,
artifact_unavailable) and, when the upstream supplied a reason, a
human-readable error_message explaining it. Prefer error for branching and
error_message for display. The distinction that matters operationally is
invalid_parameters versus upstream_error: the first means the request itself
was rejected and will fail identically on every retry (for example a seconds
value outside the model's supported range), so fix the request rather than
retrying; the second may be transient. Failed jobs are never billed.
Billing. You're charged by seconds of output video, measured when the
job completes. Jobs that never complete — failed, cancelled, or expired
before finishing — are never billed. The authoritative charge is the
cost_cents field on the job object; note that a completed (billed) job
reports status: "expired" once its 24-hour artifact window lapses, with
cost_cents still attached.
Idempotency. Send an Idempotency-Key header on POST /v1/videos so a retry
of the same request returns the original job instead of starting (and charging
for) a second one.
Retention. Completed videos are retained for 24 hours — the job object
carries an expires_at (unix seconds). After that the artifact is gone and
GET /v1/videos/{id}/content returns 410 Gone with an expired error code.
Cancel. Stop an in-flight job with POST /v1/videos/{id}/cancel (a cancel
verb, not DELETE). Cancellation is idempotent — a terminal job is returned
unchanged, and a job whose output is already committed can no longer be cancelled.
# Create (idempotent) — returns a queued job
curl https://api.deeprelay.ai/v1/videos \
-H "Authorization: Bearer deeprelay_live_…" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: 9f1c2b7e-…" \
-d '{"model":"deeprelay/seedance-1.0-lite","prompt":"ocean waves at sunset","seconds":5}'
# Poll until completed
curl https://api.deeprelay.ai/v1/videos/<video-id> -H "Authorization: Bearer deeprelay_live_…"
# Download the mp4
curl https://api.deeprelay.ai/v1/videos/<video-id>/content -H "Authorization: Bearer deeprelay_live_…" -o waves.mp4
# Cancel an in-flight job
curl -X POST https://api.deeprelay.ai/v1/videos/<video-id>/cancel -H "Authorization: Bearer deeprelay_live_…"
{ "id": "vid_…", "object": "video", "model": "deeprelay/seedance-1.0-lite",
"status": "completed", "progress": 100, "seconds": 5, "cost_cents": 18,
"created_at": 1733800000, "completed_at": 1733800240, "expires_at": 1733886640 }
The deeprelay video CLI wraps this lifecycle: create --wait polls to completion,
and download streams the mp4.
Embeddings
POST /v1/embeddings is synchronous and OpenAI-compatible. Pass input as
a single string or an array of up to 2048 strings (bounded so the response
stays under 64 MiB — about 680 inputs at 4096 dimensions; see Limits); the
response carries one vector per input, in request order, with data[].index
matching the input's position.
Find the served embedding models with GET /v1/models?modality=embedding — ids
have the usual deeprelay/{slug} form (e.g. deeprelay/qwen3-embedding-8b).
curl https://api.deeprelay.ai/v1/embeddings \
-H "Authorization: Bearer deeprelay_live_…" \
-H "Content-Type: application/json" \
-d '{"model":"deeprelay/qwen3-embedding-8b","input":["hello","world"]}'
{ "object": "list",
"data": [ { "object": "embedding", "index": 0, "embedding": [-0.019, 0.0071, …] },
{ "object": "embedding", "index": 1, "embedding": [0.004, -0.026, …] } ],
"model": "deeprelay/qwen3-embedding-8b",
"usage": { "prompt_tokens": 4, "total_tokens": 4 } }
The OpenAI SDK works unchanged:
from openai import OpenAI
client = OpenAI(base_url="https://api.deeprelay.ai/v1", api_key="deeprelay_live_…")
resp = client.embeddings.create(
model="deeprelay/qwen3-embedding-8b",
input=["first sentence", "second sentence"],
)
print(len(resp.data[0].embedding))
Billing. Embeddings are billed on input tokens only, at the model's
listed input_per_1m_tokens_cents rate, rounded up to the next whole cent
per request — so every call is billed at least 1¢. An embeddings call
produces no completion tokens, so the usage block carries prompt_tokens and
total_tokens only (no completion_tokens field): the bill is
usage.prompt_tokens × the input rate, rounded up to a whole cent. Batch
inputs into one call (up to 2048 per request) to pay the listed per-token rate
rather than the 1¢ minimum. Rollups are available from
GET /v1/usage?modality=embedding.
Limits and unsupported parameters.
- Up to 2048 inputs per request, bounded so the response stays under 64 MiB — about 680 inputs at 4096 dimensions, fewer at a larger
dimensionsvalue. An over-limit batch is rejected with a400whose message states the limit, before any processing. A single input is capped at 32 KiB, the inputs combined at 1 MiB, and the whole request body at 2 MiB + 64 KiB (413). encoding_formatsupportsfloatonly.base64is rejected with400/invalid_request_error/unsupported_parameter.inputmust be text — a single string or an array of strings. Pre-tokenized token-id arrays are not supported.dimensionsis passed through, but only models that support truncated (Matryoshka) embeddings honour it.- A chat model id on this endpoint returns
404 model_not_found: the id is valid, but not on this surface.
Audio generation
Not yet available. The serverless catalog currently serves chat, embedding, image, and video models only — there is no audio (text-to-speech or music) endpoint, and generated videos do not include an audio track. If audio matters for your workload, tell us at support@deeprelay.ai.
Tool calling
tools and tool_choice are supported on the chat models that advertise
them. The model does not execute anything — it returns the call it wants
made, and running the function and feeding the result back is your loop.
Check support first. Tool calling is per-model, not account-wide. Look for
tools in the model's supported_parameters:
curl -s https://api.deeprelay.ai/v1/models \
-H "Authorization: Bearer $DEEPRELAY_API_KEY" \
| jq -r '.data[] | select(.supported_parameters | index("tools")) | .id'
Models that support it today:
deeprelay/deepseek-v4-prodeeprelay/deepseek-v4-flashdeeprelay/deepseek-v4.1-flashdeeprelay/glm-5.2deeprelay/glm-5.3deeprelay/llama-3.3-70b-instructdeeprelay/gpt-oss-120b
Sending tools to a model that does not list it is not guaranteed to fail
cleanly — the request reaches the upstream, and what comes back is whatever
that upstream does with an unsupported parameter. Check the catalog.
Making a call
curl https://api.deeprelay.ai/v1/chat/completions \
-H "Authorization: Bearer $DEEPRELAY_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "deeprelay/deepseek-v4-pro",
"messages": [{"role": "user", "content": "What is the weather in Paris?"}],
"tools": [{
"type": "function",
"function": {
"name": "get_weather",
"description": "Current weather for a city",
"parameters": {
"type": "object",
"properties": {"city": {"type": "string"}},
"required": ["city"]
}
}
}],
"tool_choice": "auto"
}'
The reply carries finish_reason: "tool_calls" and a null content:
{
"choices": [{
"index": 0,
"message": {
"role": "assistant",
"content": null,
"tool_calls": [{
"id": "call_07db14bq3p5fkdbnv7fbsppj",
"type": "function",
"function": {"name": "get_weather", "arguments": "{\"city\": \"Paris\"}"}
}]
},
"finish_reason": "tool_calls"
}]
}
Two things about arguments that bite people:
- It is a JSON-encoded string, not an object. Parse it.
- A model can emit malformed JSON in it. Handle the parse failure rather than assuming it succeeds — this is the single most common source of a crash in a first tool-calling integration.
Closing the loop
Run the function yourself, then send the result back as a role: "tool"
message carrying the tool_call_id it answers. The assistant turn that made
the call must be replayed too, or the model has no record of having asked:
curl https://api.deeprelay.ai/v1/chat/completions \
-H "Authorization: Bearer $DEEPRELAY_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "deeprelay/deepseek-v4-pro",
"messages": [
{"role": "user", "content": "What is the weather in Paris?"},
{"role": "assistant", "content": null, "tool_calls": [{
"id": "call_07db14bq3p5fkdbnv7fbsppj",
"type": "function",
"function": {"name": "get_weather", "arguments": "{\"city\": \"Paris\"}"}
}]},
{"role": "tool", "tool_call_id": "call_07db14bq3p5fkdbnv7fbsppj",
"content": "{\"temp_c\": 14, \"condition\": \"light rain\"}"}
],
"tools": [{"type": "function", "function": {"name": "get_weather",
"parameters": {"type": "object", "properties": {"city": {"type": "string"}}}}}]
}'
Omitting tool_call_id on the tool message is a 400 from us, before the
request costs you anything.
Streaming
Tool calls stream as fragments, and a fragment is not independently
useful. The opening chunk carries id, type and function.name with empty
arguments; every chunk after it appends a slice of the arguments string:
data: {"choices":[{"delta":{"tool_calls":[{"index":0,"id":"call_pem…","type":"function","function":{"name":"get_weather","arguments":""}}]}}]}
data: {"choices":[{"delta":{"tool_calls":[{"index":0,"function":{"arguments":"{\"city\": \""}}]}}]}
data: {"choices":[{"delta":{"tool_calls":[{"index":0,"function":{"arguments":"Paris"}}]}}]}
data: {"choices":[{"delta":{"tool_calls":[{"index":0,"function":{"arguments":"\"}"}}]}}]}
data: {"choices":[{"delta":{},"finish_reason":"tool_calls"}]}
Accumulate by the index field, never by position in the tool_calls
array. A model can request several calls at once and their fragments
interleave; keying by position splices two different calls' arguments together
and produces JSON that parses cleanly and is wrong.
tool_choice
| Value | Effect |
|---|---|
"auto" | The model decides. The default when tools is present. |
"none" | Never call a tool. The only value valid with no tools array. |
"required" | Must call at least one tool. |
{"type":"function","function":{"name":"…"}} | Must call that specific function. |
Anything else is a 400 naming tool_choice.
deeprelay/glm-5.3 accepts "auto" only. Its upstream rejects the other
modes, so "none", "required" or a named function on glm-5.3 is a 400
naming tool_choice before anything is billed. Omit tool_choice or send
"auto". If you need to force a call, deeprelay/glm-5.2 supports every
mode.
From the CLI
deeprelay chat deeprelay/deepseek-v4-pro "What is the weather in Paris?" \
--tool '{"name":"get_weather","parameters":{"type":"object","properties":{"city":{"type":"string"}},"required":["city"]}}'
See deeprelay chat for --tool @file.json,
loading a whole toolset from one file, and --output json to pipe the calls
into your own dispatcher.
Coding agents depend on tool calling. To use one of these models from Cline, Roo Code, Aider, Zed or another agent, see Coding agents.
JSON mode & structured outputs
Not yet supported. Unlike tool calling, response_format is still rejected
for every model.
| Parameter | Status | What you get |
|---|---|---|
response_format: {"type":"json_object"} — JSON mode | Not yet supported | 400 unsupported_parameter, param: "response_format" |
response_format: {"type":"json_schema", …} — structured outputs | Not yet supported | 400 unsupported_parameter, param: "response_format" |
function_call (legacy) | Not supported — use tool_choice | 400 unsupported_parameter, param: "function_call" |
logprobs | Not yet supported | 400 unsupported_parameter, param: "logprobs" |
top_logprobs | Not yet supported | 400 unsupported_parameter, param: "top_logprobs" |
logit_bias | Not yet supported | 400 unsupported_parameter, param: "logit_bias" |
Sending any of them gets you this exact body, with param naming the field
that was rejected:
{ "error": { "message": "Parameter \"response_format\" is not supported. See https://deeprelay.ai/docs/api/inference#openai-compatibility for what the chat endpoint accepts.", "type": "invalid_request_error", "param": "response_format", "code": "unsupported_parameter" } }
If more than one is present, only the first is named; the order the request is
checked in is logprobs, top_logprobs, function_call, response_format,
logit_bias.
If you need guaranteed-shape output today, a tool with the schema you want as
its parameters and tool_choice forcing that function gets you most of the
way there on a tool-capable model — the model fills your schema, and you read
it out of arguments instead of out of the message text.
Workaround: prompted JSON + client-side validation
Ask for JSON in the prompt, then parse and validate it yourself, retrying when
the model returns something that does not parse or does not match your schema.
Be explicit that this is best-effort: without response_format the model is
not constrained by the sampler, so a small fraction of responses will need the
retry.
curl https://api.deeprelay.ai/v1/chat/completions \
-H "Authorization: Bearer deeprelay_live_…" \
-H "Content-Type: application/json" \
-d '{
"model": "deeprelay/llama-3.3-70b-instruct",
"temperature": 0,
"messages": [
{"role": "system", "content": "Reply with a single JSON object and nothing else. No prose, no markdown fences. Schema: {\"capital\": string}."},
{"role": "user", "content": "What is the capital of France?"}
]
}'
import json
from jsonschema import ValidationError, validate
from openai import OpenAI
client = OpenAI(base_url="https://api.deeprelay.ai/v1", api_key="deeprelay_live_…")
SCHEMA = {
"type": "object",
"properties": {"capital": {"type": "string"}},
"required": ["capital"],
"additionalProperties": False,
}
SYSTEM = (
"Reply with a single JSON object and nothing else. "
"No prose, no markdown fences. "
f"It must validate against this JSON Schema: {json.dumps(SCHEMA)}"
)
def structured(question: str, attempts: int = 3) -> dict:
"""Best-effort structured output: prompt for JSON, parse, validate, retry."""
last_error = None
for _ in range(attempts):
resp = client.chat.completions.create(
model="deeprelay/llama-3.3-70b-instruct",
temperature=0,
messages=[
{"role": "system", "content": SYSTEM},
{"role": "user", "content": question},
],
)
text = resp.choices[0].message.content.strip()
# Models sometimes wrap JSON in a fence despite the instruction.
if text.startswith("```"):
text = text.strip("`").removeprefix("json").strip()
try:
data = json.loads(text)
validate(instance=data, schema=SCHEMA)
return data
except (json.JSONDecodeError, ValidationError) as err:
last_error = err
raise RuntimeError(f"no valid JSON after {attempts} attempts: {last_error}")
print(structured("What is the capital of France?"))
Pick any served chat model from GET /v1/models?modality=chat —
deeprelay/llama-3.3-70b-instruct above is one of them. temperature: 0 and a
short, explicit schema in the system prompt materially cut the retry rate;
jsonschema can be swapped for pydantic if you already model your responses
there.
Roadmap
tools and tool_choice have shipped, streaming delta.tool_calls
included — see Tool calling above.
response_format has not. It stays rejected rather than silently ignored, so
you never get an unconstrained answer back from a request that asked for a
constrained one. If it blocks you, tell us at support@deeprelay.ai — demand
ordering is how this gets scheduled.
Errors
Errors use the OpenAI envelope, so SDK error handling works as-is (the /v1/messages endpoints use the Anthropic envelope instead; see Anthropic Messages API):
{ "error": { "message": "…", "type": "invalid_request_error", "code": "model_not_found" } }
Error responses also carry an X-Deeprelay-Error-Code header set to the same
code (for example model_not_found, rate_limit_exceeded,
insufficient_balance). The header uses one set of codes across every endpoint,
so you can branch on or count errors without parsing the body. (An error that
arrives mid-stream as an SSE error event carries its code in the event body
only.)
Model availability. The serverless catalog is live, so a model id that worked
earlier can stop being served. When that happens you get a model_not_found
(404) with a "Try another model" message — re-fetch GET /v1/models for the
current list and pick another. On the streaming endpoint this arrives as the
final SSE error event before [DONE].
Concurrent-stream cap. Each API key may have a limited number of inference
requests open at once. If you exceed it you get a 429 with code
stream_limit_exceeded:
{ "error": { "message": "Too many concurrent streams open for this API key — close existing streams or retry shortly.", "type": "rate_limit_error", "code": "stream_limit_exceeded" } }
Close any streams you are no longer reading, or retry shortly — the slot frees
as soon as an in-flight request finishes. This is separate from the per-second
rate limit (rate_limit_exceeded), which bounds request rate rather than the
number of simultaneously-open requests.
Upstream temporarily unavailable. If a model's serving capacity is degraded,
requests may return a 503 with code upstream_unavailable:
{ "error": { "message": "Upstream service temporarily unavailable.", "type": "api_error", "code": "upstream_unavailable" } }
This is returned promptly (no long wait) while capacity recovers — retry after a
short interval. The platform automatically routes around a degraded endpoint and
resumes serving it once it recovers, so a 503 here is transient.