deeprelayDocs
CH·GGuides

Serverless Inference API

The OpenAI-compatible /v1 API in full: chat, embeddings, image and video, the model catalog, tool calling, the Anthropic Messages endpoint, plan coverage, and errors.

deeprelay's serverless inference is an OpenAI-compatible HTTP API: point any OpenAI SDK (or curl, or the deeprelay CLI) at it and call chat completions, embeddings, image generation, video generation, and the model catalog. You pay per token (or per image / per video-second) — no instances to manage.

Base URL & authentication

Base URLhttps://api.deeprelay.ai/v1
AuthAuthorization: Bearer deeprelay_live_…

Get a key with deeprelay login (it stores a deeprelay_live_… key in ~/.config/deeprelay/credentials.json) or mint one in the dashboard. Scopes: serverless:read lists models, serverless:write runs chat/embeddings/image/video, billing:read reads usage; a full_access key covers all three.

Every key you mint belongs to this production host. api.demo.deeprelay.ai, which appears in some older examples, is our internal staging environment with its own key store, so a production key gets a 401 there. /v1 is the API version and the only one there is; the OpenAPI document's 1.0.0 is the revision of that /v1 contract.

Trial keys

If someone at deeprelay sent you a trial key, it is a normal deeprelay_live_… API key with trial credit already loaded — no account or sign-in needed. Use it exactly like any other key against the base URL above:

export DEEPRELAY_API_KEY=deeprelay_live_…   # the key you were sent

curl https://api.deeprelay.ai/v1/chat/completions \
  -H "Authorization: Bearer $DEEPRELAY_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"deeprelay/deepseek-v4-pro","messages":[{"role":"user","content":"Hello!"}]}'

Or with an OpenAI SDK, set base_url="https://api.deeprelay.ai/v1" and api_key to the trial key. GET /v1/models lists every model and its per-token price; the DeepSeek, GLM and Kimi families are the ones the trial is meant for.

What a trial key can and cannot do:

  • Scope: serverless inference (serverless:write) plus reading its own balance (billing:read). It cannot launch GPU instances or mint other keys.
  • Credit is the cap. GET /v1/billing/balance shows what is left. When the balance reaches zero, requests return 402 insufficient_balance.
  • It expires on the date in the message you received (30 days by default). After that, inference calls (/v1/chat/completions, /v1/models and the other OpenAI-compatible routes) return 401 with the OpenAI-style error code unauthorized; GET /v1/billing/balance returns the RFC 7807 problem code invalid_api_key.
  • Rate limit is 10 requests/second, enough for trying models by hand.
  • To keep going after the trial, sign up at deeprelay.ai and mint your own key — trial credit does not transfer.

Plan coverage

Models the flat subscription tier covers carry "plan_covered": true in GET /v1/models (and a PLAN/Plan marker in deeprelay models list / deeprelay models get) — a subscriber's requests to those models draw on the plan's monthly allowance first; everything else bills pay-as-you-go. GET /v1/billing/subscription (CLI: deeprelay billing subscription) reports how much of that allowance is left.

Being subscribed does not make every model free, and the response to a request for an uncovered model does not mention that it was charged. So ask first: GET /v1/inference/preflight?model= returns ok / warn / block for that exact model, computed from the same gates that will judge the request — plan coverage, entitlement, remaining quota, credit and spending caps. The deeprelay CLI runs it automatically before chat, embeddings, image and video create, warning when a call will bill pay-as-you-go and refusing when it would be denied; deeprelay preflight asks it directly, and --no-preflight turns the automatic check off.

Tiers — serverless (default) and economy

A model can have two tiers:

  • serverless (default) — low latency, no cold start. Use the bare model id, e.g. deeprelay/deepseek-v4-flash.
  • economy — cheaper, compute-priced, may cold-start ~30–60s. Opt in with the :economy suffix, e.g. deeprelay/:economy.

Economy exists only where GET /v1/models lists the suffixed id as its own entry. Check the live catalog before relying on it: a :economy id the catalog does not list returns model_not_found, and there are times when no model has an economy tier at all.

deeprelay models list shows both tiers as separate rows with their own pricing.

The serverless catalog is live

The serverless catalog is built from the models our upstream partners are actually serving warm right now — not a hand-maintained list. It is refreshed roughly hourly and probe-verified, so GET /v1/models always reflects what you can call this minute. Each entry carries name, author (the model's creator — e.g. Meta, Qwen, Black Forest Labs), category (mirrors modality: chat, embeddings, image, or video), and pricing (the cheaper option when a model is served by more than one partner). Chat models are priced per 1M tokens (input_per_1m_tokens_cents / output_per_1m_tokens_cents), plus an optional discounted rate for the cache-hit part of the prompt (cached_input_per_1m_tokens_microcents — see Prompt caching below). Embedding models are priced per 1M input tokens only (input_per_1m_tokens_cents); they emit no output tokens, so output_per_1m_tokens_cents is absent. Image models are priced per output megapixel when the partner meters that way (per_mpxl_microcents, in micro-cents — e.g. 270000 = $0.0027/Mpxl), otherwise with a representative per-image "from" price (per_image_microcents). Video models are billed per second of output video (per_video_second_microcents); the per-clip per_video_microcents value is the representative "from" price for the model's default clip. The set of authors we surface grows over time.

Because the catalog is live, a model can occasionally drop out between refreshes. If you call one that just became unavailable you'll get a clear model_not_found error asking you to try another — see Errors below.

Using the deeprelay CLI

# Discover models
deeprelay models list --modality chat
deeprelay models get deeprelay/deepseek-v4-flash

# Chat (streams tokens)
deeprelay chat deeprelay/deepseek-v4-flash "Explain WireGuard in one sentence"

# Embeddings (one vector per input; -o json for the full vectors)
deeprelay embeddings deeprelay/qwen3-embedding-8b "first sentence" "second sentence"

# Image generation (writes out.png)
deeprelay image "a watercolor fox in a misty forest" --model deeprelay/flux.2-dev

# Video generation (async: submit, poll, download)
deeprelay video create --model deeprelay/seedance-1.0-lite --prompt "ocean waves at sunset" --wait 10m
deeprelay video download <video-id> --out waves.mp4

# Usage / spend
deeprelay usage --modality chat --bucket month

# On the flat tier: status plus how much quota is left
deeprelay billing subscription

# Will this model draw on the plan, or on credit?
deeprelay preflight deeprelay/deepseek-v4-flash

# Start or cancel a subscription (each prints a link to open)
deeprelay billing subscribe
deeprelay billing cancel

# Refer a friend: you both get $5 in credit once they spend $5
deeprelay referral

See the per-command reference: chat, embeddings, image, video, models, usage, preflight, billing subscription, billing subscribe, billing cancel, referral (see the refer-a-friend guide).

OpenAI compatibility

The chat request and response shapes, streaming, tool calling and the error envelope match OpenAI's. Not every OpenAI chat parameter is carried, though. Each one is forwarded to the model, rejected with a 400 before anything is billed, or ignored (accepted and dropped):

ParameterHandling
model, messagesForwarded. content may be a string, null, or an array of text parts ([{"type":"text","text":"…"}], joined with newlines). Roles are system, user, assistant and tool; developer is treated as system.
temperature, top_p, max_tokens, stop, n, seed, userForwarded.
max_completion_tokensForwarded as max_tokens. If you send both, max_tokens wins.
stream, stream_optionsForwarded.
tools, tool_choiceForwarded, on models that list tools in supported_parameters (see Tool calling). finish_reason is tool_calls when the model makes a call.
response_format, logprobs, top_logprobs, logit_bias, function_callRejected: 400 unsupported_parameter (see JSON mode).
image_url, input_audio, file content partsRejected: 400 unsupported_content_part, param: "messages". No model in the catalog takes image, audio or file input.
reasoning_effort, presence_penalty, frequency_penalty, parallel_tool_calls, anything elseIgnored: accepted without error and not sent to the model.

Ignored means not applied. reasoning_effort is the one most likely to matter: a reasoning model thinks at its own default effort whatever you send, and its reasoning still streams as reasoning deltas. Don't rely on an ignored parameter for correctness.

Using the OpenAI SDK

Because the API is OpenAI-compatible, the official SDKs work unchanged for chat (within the table above). Just override base_url and api_key:

from openai import OpenAI

client = OpenAI(
    base_url="https://api.deeprelay.ai/v1",
    api_key="deeprelay_live_…",
)

# Chat (streaming)
stream = client.chat.completions.create(
    model="deeprelay/deepseek-v4-flash",
    messages=[{"role": "user", "content": "Explain WireGuard in one sentence"}],
    stream=True,
)
for chunk in stream:
    print(chunk.choices[0].delta.content or "", end="")

Coding agents

The same base URL and key work in coding agents that accept a custom OpenAI-compatible provider, such as Cline, Kilo Code, Roo Code, OpenCode, Aider, Continue, Zed and Cursor. Coding agents has the config for each and the models to pick for tool calling.

Claude Code uses the Anthropic Messages API instead. Run deeprelay setup claude-code, and see Claude Code and Anthropic Messages API below.

Using curl

# List models
curl https://api.deeprelay.ai/v1/models \
  -H "Authorization: Bearer deeprelay_live_…"

# Chat completion (non-streaming)
curl https://api.deeprelay.ai/v1/chat/completions \
  -H "Authorization: Bearer deeprelay_live_…" \
  -H "Content-Type: application/json" \
  -d '{"model":"deeprelay/deepseek-v4-flash","messages":[{"role":"user","content":"hi"}]}'

Anthropic Messages API

deeprelay also serves the Anthropic Messages API at POST /v1/messages, so Claude Code and the Anthropic SDKs can use deeprelay models. For Claude Code, run deeprelay setup claude-code. Claude Code is the full guide. The Anthropic SDKs take the base URL without /v1 (https://api.deeprelay.ai).

curl https://api.deeprelay.ai/v1/messages \
  -H "x-api-key: deeprelay_live_…" \
  -H "anthropic-version: 2023-06-01" \
  -H "Content-Type: application/json" \
  -d '{"model":"deeprelay/deepseek-v4-pro","max_tokens":256,"messages":[{"role":"user","content":"hi"}]}'
  • Auth: send the key as x-api-key or as Authorization: Bearer.
  • Models: use deeprelay model IDs. A claude-* model ID is rejected with 400 (claude_model_not_served) and a message telling you to run deeprelay setup claude-code. deeprelay doesn't alias Claude names to other models.
  • Supported: text, tools, streaming (with ping keep-alives), thinking and prompt caching. cache_control markers are accepted, and caching happens automatically where the model supports it.
  • Not supported, each rejected with a 400:
    • server-side tools such as web search, web fetch and code execution (server_tool_unsupported). Use an MCP server in the client instead;
    • mcp_servers and container (server_side_feature_unsupported);
    • document and PDF blocks;
    • a user image in the newest user message on a model without vision support (unsupported_capability; the message names the models that accept images). Today only deeprelay/deepseek-v4.1-flash accepts images on this endpoint.
  • Images that are not rejected: on a model without vision, an image in an earlier turn or inside a tool_result is replaced by a text note saying an image was omitted. On a vision model, an image inside a tool_result is moved into user content right after the tool result, because most models take tool results as text only. The response header x-deeprelay-ignored-params then lists image:omitted and/or image:moved.
  • Stop sequences: stop_sequences are honoured, but when one ends the response the result says stop_reason: "end_turn" with stop_sequence: null, because the models behind this endpoint do not report which sequence matched.
  • Token counting: POST /v1/messages/count_tokens returns an estimate, marked with the header x-deeprelay-token-count: estimated. It is never billed.
  • Errors use the Anthropic envelope, {"type":"error","error":{"type":"…","message":"…"},"request_id":"…"}. The deeprelay code rides in the X-Deeprelay-Error-Code header, from the same set of codes as the rest of the API (see Errors).
  • Billing is the same as a chat completion on the same model. Rejected requests are not billed.

Endpoints

MethodPathScopePurpose
GET/v1/models (?modality=)serverless:readList the model catalog
GET/v1/models/{id}serverless:readOne model's record
POST/v1/chat/completionsserverless:writeChat completion (supports stream:true)
POST/v1/messagesserverless:writeAnthropic Messages API, for Claude Code (supports stream:true); see Anthropic Messages API
POST/v1/messages/count_tokensserverless:readEstimated input-token count for a Messages request (local estimate, not billed)
POST/v1/embeddingsserverless:writeText embeddings (OpenAI-compatible, synchronous)
POST/v1/images/generationsserverless:writeSynchronous image generation (b64_json)
POST/v1/videosserverless:writeCreate an async video job
GET/v1/videos (?after= / ?limit=)serverless:readList your video jobs
GET/v1/videos/{id}serverless:readOne video job (poll for status)
GET/v1/videos/{id}/contentserverless:readDownload the finished mp4
POST/v1/videos/{id}/cancelserverless:writeCancel an in-flight video job
GET/v1/usage (?modality= / ?model=)billing:readToken/cost rollups
GET/v1/billing/subscriptionbilling:readFlat-tier status, plan + price, quota usage (200 with subscribed:false when not subscribed)
POST/v1/billing/subscription/checkoutbilling:write (org admin)Mint a Stripe Checkout URL to subscribe (409 if already subscribed)
POST/v1/billing/subscription/portalbilling:write (org admin)Mint a Stripe Billing Portal URL to cancel, resume or change card
GET/v1/inference/preflight (?model=)serverless:readWould this model request be served, and at whose expense

Reading your usage and spend

GET /v1/usage returns your inference spend in time buckets (bucket=hour, day, week or month; default day) over start/end (RFC 3339; default the last 30 days). Pass modality and/or model — those filters are what select serverless inference usage. Without either, the endpoint reads GPU instance usage, which is always empty today.

# Chat spend per day for the last 30 days
curl "https://api.deeprelay.ai/v1/usage?modality=chat" \
  -H "Authorization: Bearer $DEEPRELAY_API_KEY"

# One model, hourly, over an explicit window
curl "https://api.deeprelay.ai/v1/usage?model=<model-id>&bucket=hour&start=2026-09-01T00:00:00Z&end=2026-09-02T00:00:00Z" \
  -H "Authorization: Bearer $DEEPRELAY_API_KEY"

modality is one of chat, image, video or embedding; model is an exact model id from GET /v1/models. Each row covers one (bucket, modality, model) and carries bucket_start, modality, model, prompt_tokens, completion_tokens, image_count and cost_cents (plus gpu_seconds, always 0 on inference rows). Rows come newest bucket first, paginated with limit/cursor like every other list. The CLI equivalent is deeprelay usage --modality chat (or --model ), and the SDKs expose both filters on listUsage / list_usage.

Prompt caching (cached input tokens)

When an upstream partner serves part of your prompt from its prompt cache, we pass the split straight through and bill the cached part at a lower rate.

In the response. A chat completion's usage block carries prompt_tokens_details.cached_tokens whenever the upstream reported a non-zero cache hit — the same field name and shape the OpenAI SDKs already read, on both streaming and non-streaming calls:

"usage": {
  "prompt_tokens": 4096,
  "completion_tokens": 128,
  "total_tokens": 4224,
  "prompt_tokens_details": { "cached_tokens": 3840 }
}

cached_tokens is a subset of prompt_tokens, never an addition: prompt_tokens remains the full prompt. When there were no cache hits the prompt_tokens_details object is omitted entirely.

In the catalog. A model with a cached-input rate exposes it in pricing as cached_input_per_1m_tokens_microcents (micro-cents per 1M tokens, the exact figure — cached rates are routinely sub-cent) alongside a rounded cached_input_per_1m_tokens_cents. Treat the micro-cents field as the presence signal: it is absent only when the model has no cached rate (no discount — every prompt token bills at input_per_1m_tokens_cents), while the rounded cents field is also omitted whenever the exact rate rounds below half a cent, which for cached rates is the routine case.

How a chat call is billed.

cost = (prompt_tokens - cached_tokens) × input_rate
     +  cached_tokens                  × cached_input_rate
     +  completion_tokens              × output_rate

rounded up to the next whole cent per request, with all rates per 1M tokens. cached_tokens is clamped to prompt_tokens, so an over-reporting upstream can never produce a negative cache-miss count. Prompt caching is a partner-side optimisation — we don't control which prefixes get cached, and a call that hits no cache simply bills exactly as it did before.

Where the cached rate comes from. A model's cached-input rate is taken from the serving partner's published cache-hit price (the same listing its input and output rates come from) and carries the same retail markup, or from the pricing policy set for the model when one is enrolled. It is visible everywhere the model's price is: GET /v1/models, deeprelay models list / deeprelay models get, the public model pricing pages and the console's model catalog. A model that shows no cached rate has none — its partner publishes no cache-hit price — and bills every prompt token at the input rate.

Peak and off-peak pricing

A few models are priced by time of day: their upstream charges more inside daily peak hours. For those models pricing carries the whole picture:

"pricing": {
  "currency": "usd",
  "input_per_1m_tokens_cents": 30,
  "output_per_1m_tokens_cents": 89,
  "cached_input_per_1m_tokens_microcents": 114000,
  "period": "peak",
  "peak_windows_utc": [
    {"start": "01:00", "end": "04:00"},
    {"start": "06:00", "end": "10:00"}
  ],
  "peak":     {"input_per_1m_tokens_microcents": 30180000, "cached_input_per_1m_tokens_microcents": 114000, "output_per_1m_tokens_microcents": 88620000, "...": "..."},
  "off_peak": {"input_per_1m_tokens_microcents": 15090000, "cached_input_per_1m_tokens_microcents": 57000,  "output_per_1m_tokens_microcents": 44310000, "...": "..."}
}
  • The flat fields (input_per_1m_tokens_cents, output_per_1m_tokens_cents, cached_input_per_1m_tokens_microcents) are the rates in effect at the moment you read the catalog, and period says which side of the schedule that is (peak or off_peak). A client that reads only the flat fields always sees the price a request sent right now is billed at.
  • peak_windows_utc is the daily schedule in UTC — start inclusive, end exclusive, "24:00" allowed as an end, and a range whose end sorts before its start wraps midnight. peak and off_peak carry both full rate triples so you can price a call you will make later.
  • A request is billed at the rate in effect when it arrives (its UTC arrival minute is checked against the windows). A long streamed response that straddles a boundary is priced entirely at the arrival rate.
  • Models without time-of-day pricing omit all four fields; their flat rates apply around the clock.

deeprelay models get shows the period, the windows and both triples for such a model.

Balance checks and usage rollups. The pre-flight balance check assumes zero cache hits, so it stays conservative: a cached call can only cost less than the amount reserved. GET /v1/usage and the Stripe input-token meter keep counting the full prompt_tokens — the cache discount shows up in the cost, not in the token counts.

Image generation

POST /v1/images/generations is synchronous and OpenAI-compatible. v1 returns images as base64 only — data[].b64_json carries the bytes, and response_format:"url" is rejected with an invalid_request_error. The response adds a usage.image_count field (the metered unit).

curl https://api.deeprelay.ai/v1/images/generations \
  -H "Authorization: Bearer deeprelay_live_…" \
  -H "Content-Type: application/json" \
  -d '{"model":"deeprelay/flux.2-dev","prompt":"a watercolor fox","n":1,"size":"1024x1024"}'
{ "created": 1733800000, "data": [{ "b64_json": "iVBORw0KGgo…" }], "usage": { "image_count": 1 } }

The deeprelay image CLI decodes the base64 and writes the file(s) for you:

deeprelay image "a watercolor fox in a misty forest" --model deeprelay/flux.2-dev --out fox.png

Video generation

Video generation is asynchronous. POST /v1/videos returns a job in the queued state immediately; poll GET /v1/videos/{id} until the status reaches a terminal value, then download the mp4 from GET /v1/videos/{id}/content.

Lifecycle. A job moves queued → in_progress → completed (or failed / cancelled / expired). While in_progress the job carries a progress value (0–100); some models report no granular progress and stay at 0 until they complete.

Parameters. Each model declares the request parameters it accepts in its catalog entry (parameters on GET /v1/models) — typically seconds (clip length; bounds and default come from the descriptor). Parameters a model does not declare are rejected with an invalid_request_error rather than silently ignored: live video models generate at their default output resolution, so most do not accept size. Omitting seconds uses the model's default clip length.

Failures. A failed job carries a canonical code in error (invalid_parameters, content_policy, upstream_error, artifact_unavailable) and, when the upstream supplied a reason, a human-readable error_message explaining it. Prefer error for branching and error_message for display. The distinction that matters operationally is invalid_parameters versus upstream_error: the first means the request itself was rejected and will fail identically on every retry (for example a seconds value outside the model's supported range), so fix the request rather than retrying; the second may be transient. Failed jobs are never billed.

Billing. You're charged by seconds of output video, measured when the job completes. Jobs that never complete — failed, cancelled, or expired before finishing — are never billed. The authoritative charge is the cost_cents field on the job object; note that a completed (billed) job reports status: "expired" once its 24-hour artifact window lapses, with cost_cents still attached.

Idempotency. Send an Idempotency-Key header on POST /v1/videos so a retry of the same request returns the original job instead of starting (and charging for) a second one.

Retention. Completed videos are retained for 24 hours — the job object carries an expires_at (unix seconds). After that the artifact is gone and GET /v1/videos/{id}/content returns 410 Gone with an expired error code.

Cancel. Stop an in-flight job with POST /v1/videos/{id}/cancel (a cancel verb, not DELETE). Cancellation is idempotent — a terminal job is returned unchanged, and a job whose output is already committed can no longer be cancelled.

# Create (idempotent) — returns a queued job
curl https://api.deeprelay.ai/v1/videos \
  -H "Authorization: Bearer deeprelay_live_…" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: 9f1c2b7e-…" \
  -d '{"model":"deeprelay/seedance-1.0-lite","prompt":"ocean waves at sunset","seconds":5}'

# Poll until completed
curl https://api.deeprelay.ai/v1/videos/<video-id> -H "Authorization: Bearer deeprelay_live_…"

# Download the mp4
curl https://api.deeprelay.ai/v1/videos/<video-id>/content -H "Authorization: Bearer deeprelay_live_…" -o waves.mp4

# Cancel an in-flight job
curl -X POST https://api.deeprelay.ai/v1/videos/<video-id>/cancel -H "Authorization: Bearer deeprelay_live_…"
{ "id": "vid_…", "object": "video", "model": "deeprelay/seedance-1.0-lite",
  "status": "completed", "progress": 100, "seconds": 5, "cost_cents": 18,
  "created_at": 1733800000, "completed_at": 1733800240, "expires_at": 1733886640 }

The deeprelay video CLI wraps this lifecycle: create --wait polls to completion, and download streams the mp4.

Embeddings

POST /v1/embeddings is synchronous and OpenAI-compatible. Pass input as a single string or an array of up to 2048 strings (bounded so the response stays under 64 MiB — about 680 inputs at 4096 dimensions; see Limits); the response carries one vector per input, in request order, with data[].index matching the input's position.

Find the served embedding models with GET /v1/models?modality=embedding — ids have the usual deeprelay/{slug} form (e.g. deeprelay/qwen3-embedding-8b).

curl https://api.deeprelay.ai/v1/embeddings \
  -H "Authorization: Bearer deeprelay_live_…" \
  -H "Content-Type: application/json" \
  -d '{"model":"deeprelay/qwen3-embedding-8b","input":["hello","world"]}'
{ "object": "list",
  "data": [ { "object": "embedding", "index": 0, "embedding": [-0.019, 0.0071, …] },
            { "object": "embedding", "index": 1, "embedding": [0.004, -0.026, …] } ],
  "model": "deeprelay/qwen3-embedding-8b",
  "usage": { "prompt_tokens": 4, "total_tokens": 4 } }

The OpenAI SDK works unchanged:

from openai import OpenAI
client = OpenAI(base_url="https://api.deeprelay.ai/v1", api_key="deeprelay_live_…")

resp = client.embeddings.create(
    model="deeprelay/qwen3-embedding-8b",
    input=["first sentence", "second sentence"],
)
print(len(resp.data[0].embedding))

Billing. Embeddings are billed on input tokens only, at the model's listed input_per_1m_tokens_cents rate, rounded up to the next whole cent per request — so every call is billed at least 1¢. An embeddings call produces no completion tokens, so the usage block carries prompt_tokens and total_tokens only (no completion_tokens field): the bill is usage.prompt_tokens × the input rate, rounded up to a whole cent. Batch inputs into one call (up to 2048 per request) to pay the listed per-token rate rather than the 1¢ minimum. Rollups are available from GET /v1/usage?modality=embedding.

Limits and unsupported parameters.

  • Up to 2048 inputs per request, bounded so the response stays under 64 MiB — about 680 inputs at 4096 dimensions, fewer at a larger dimensions value. An over-limit batch is rejected with a 400 whose message states the limit, before any processing. A single input is capped at 32 KiB, the inputs combined at 1 MiB, and the whole request body at 2 MiB + 64 KiB (413).
  • encoding_format supports float only. base64 is rejected with 400 / invalid_request_error / unsupported_parameter.
  • input must be text — a single string or an array of strings. Pre-tokenized token-id arrays are not supported.
  • dimensions is passed through, but only models that support truncated (Matryoshka) embeddings honour it.
  • A chat model id on this endpoint returns 404 model_not_found: the id is valid, but not on this surface.

Audio generation

Not yet available. The serverless catalog currently serves chat, embedding, image, and video models only — there is no audio (text-to-speech or music) endpoint, and generated videos do not include an audio track. If audio matters for your workload, tell us at support@deeprelay.ai.

Tool calling

tools and tool_choice are supported on the chat models that advertise them. The model does not execute anything — it returns the call it wants made, and running the function and feeding the result back is your loop.

Check support first. Tool calling is per-model, not account-wide. Look for tools in the model's supported_parameters:

curl -s https://api.deeprelay.ai/v1/models \
  -H "Authorization: Bearer $DEEPRELAY_API_KEY" \
| jq -r '.data[] | select(.supported_parameters | index("tools")) | .id'

Models that support it today:

  • deeprelay/deepseek-v4-pro
  • deeprelay/deepseek-v4-flash
  • deeprelay/deepseek-v4.1-flash
  • deeprelay/glm-5.2
  • deeprelay/glm-5.3
  • deeprelay/llama-3.3-70b-instruct
  • deeprelay/gpt-oss-120b

Sending tools to a model that does not list it is not guaranteed to fail cleanly — the request reaches the upstream, and what comes back is whatever that upstream does with an unsupported parameter. Check the catalog.

Making a call

curl https://api.deeprelay.ai/v1/chat/completions \
  -H "Authorization: Bearer $DEEPRELAY_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "deeprelay/deepseek-v4-pro",
    "messages": [{"role": "user", "content": "What is the weather in Paris?"}],
    "tools": [{
      "type": "function",
      "function": {
        "name": "get_weather",
        "description": "Current weather for a city",
        "parameters": {
          "type": "object",
          "properties": {"city": {"type": "string"}},
          "required": ["city"]
        }
      }
    }],
    "tool_choice": "auto"
  }'

The reply carries finish_reason: "tool_calls" and a null content:

{
  "choices": [{
    "index": 0,
    "message": {
      "role": "assistant",
      "content": null,
      "tool_calls": [{
        "id": "call_07db14bq3p5fkdbnv7fbsppj",
        "type": "function",
        "function": {"name": "get_weather", "arguments": "{\"city\": \"Paris\"}"}
      }]
    },
    "finish_reason": "tool_calls"
  }]
}

Two things about arguments that bite people:

  • It is a JSON-encoded string, not an object. Parse it.
  • A model can emit malformed JSON in it. Handle the parse failure rather than assuming it succeeds — this is the single most common source of a crash in a first tool-calling integration.

Closing the loop

Run the function yourself, then send the result back as a role: "tool" message carrying the tool_call_id it answers. The assistant turn that made the call must be replayed too, or the model has no record of having asked:

curl https://api.deeprelay.ai/v1/chat/completions \
  -H "Authorization: Bearer $DEEPRELAY_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "deeprelay/deepseek-v4-pro",
    "messages": [
      {"role": "user", "content": "What is the weather in Paris?"},
      {"role": "assistant", "content": null, "tool_calls": [{
        "id": "call_07db14bq3p5fkdbnv7fbsppj",
        "type": "function",
        "function": {"name": "get_weather", "arguments": "{\"city\": \"Paris\"}"}
      }]},
      {"role": "tool", "tool_call_id": "call_07db14bq3p5fkdbnv7fbsppj",
       "content": "{\"temp_c\": 14, \"condition\": \"light rain\"}"}
    ],
    "tools": [{"type": "function", "function": {"name": "get_weather",
      "parameters": {"type": "object", "properties": {"city": {"type": "string"}}}}}]
  }'

Omitting tool_call_id on the tool message is a 400 from us, before the request costs you anything.

Streaming

Tool calls stream as fragments, and a fragment is not independently useful. The opening chunk carries id, type and function.name with empty arguments; every chunk after it appends a slice of the arguments string:

data: {"choices":[{"delta":{"tool_calls":[{"index":0,"id":"call_pem…","type":"function","function":{"name":"get_weather","arguments":""}}]}}]}
data: {"choices":[{"delta":{"tool_calls":[{"index":0,"function":{"arguments":"{\"city\": \""}}]}}]}
data: {"choices":[{"delta":{"tool_calls":[{"index":0,"function":{"arguments":"Paris"}}]}}]}
data: {"choices":[{"delta":{"tool_calls":[{"index":0,"function":{"arguments":"\"}"}}]}}]}
data: {"choices":[{"delta":{},"finish_reason":"tool_calls"}]}

Accumulate by the index field, never by position in the tool_calls array. A model can request several calls at once and their fragments interleave; keying by position splices two different calls' arguments together and produces JSON that parses cleanly and is wrong.

tool_choice

ValueEffect
"auto"The model decides. The default when tools is present.
"none"Never call a tool. The only value valid with no tools array.
"required"Must call at least one tool.
{"type":"function","function":{"name":"…"}}Must call that specific function.

Anything else is a 400 naming tool_choice.

deeprelay/glm-5.3 accepts "auto" only. Its upstream rejects the other modes, so "none", "required" or a named function on glm-5.3 is a 400 naming tool_choice before anything is billed. Omit tool_choice or send "auto". If you need to force a call, deeprelay/glm-5.2 supports every mode.

From the CLI

deeprelay chat deeprelay/deepseek-v4-pro "What is the weather in Paris?" \
  --tool '{"name":"get_weather","parameters":{"type":"object","properties":{"city":{"type":"string"}},"required":["city"]}}'

See deeprelay chat for --tool @file.json, loading a whole toolset from one file, and --output json to pipe the calls into your own dispatcher.

Coding agents depend on tool calling. To use one of these models from Cline, Roo Code, Aider, Zed or another agent, see Coding agents.

JSON mode & structured outputs

Not yet supported. Unlike tool calling, response_format is still rejected for every model.

ParameterStatusWhat you get
response_format: {"type":"json_object"} — JSON modeNot yet supported400 unsupported_parameter, param: "response_format"
response_format: {"type":"json_schema", …} — structured outputsNot yet supported400 unsupported_parameter, param: "response_format"
function_call (legacy)Not supported — use tool_choice400 unsupported_parameter, param: "function_call"
logprobsNot yet supported400 unsupported_parameter, param: "logprobs"
top_logprobsNot yet supported400 unsupported_parameter, param: "top_logprobs"
logit_biasNot yet supported400 unsupported_parameter, param: "logit_bias"

Sending any of them gets you this exact body, with param naming the field that was rejected:

{ "error": { "message": "Parameter \"response_format\" is not supported. See https://deeprelay.ai/docs/api/inference#openai-compatibility for what the chat endpoint accepts.", "type": "invalid_request_error", "param": "response_format", "code": "unsupported_parameter" } }

If more than one is present, only the first is named; the order the request is checked in is logprobs, top_logprobs, function_call, response_format, logit_bias.

If you need guaranteed-shape output today, a tool with the schema you want as its parameters and tool_choice forcing that function gets you most of the way there on a tool-capable model — the model fills your schema, and you read it out of arguments instead of out of the message text.

Workaround: prompted JSON + client-side validation

Ask for JSON in the prompt, then parse and validate it yourself, retrying when the model returns something that does not parse or does not match your schema. Be explicit that this is best-effort: without response_format the model is not constrained by the sampler, so a small fraction of responses will need the retry.

curl https://api.deeprelay.ai/v1/chat/completions \
  -H "Authorization: Bearer deeprelay_live_…" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "deeprelay/llama-3.3-70b-instruct",
    "temperature": 0,
    "messages": [
      {"role": "system", "content": "Reply with a single JSON object and nothing else. No prose, no markdown fences. Schema: {\"capital\": string}."},
      {"role": "user", "content": "What is the capital of France?"}
    ]
  }'
import json

from jsonschema import ValidationError, validate
from openai import OpenAI

client = OpenAI(base_url="https://api.deeprelay.ai/v1", api_key="deeprelay_live_…")

SCHEMA = {
    "type": "object",
    "properties": {"capital": {"type": "string"}},
    "required": ["capital"],
    "additionalProperties": False,
}

SYSTEM = (
    "Reply with a single JSON object and nothing else. "
    "No prose, no markdown fences. "
    f"It must validate against this JSON Schema: {json.dumps(SCHEMA)}"
)


def structured(question: str, attempts: int = 3) -> dict:
    """Best-effort structured output: prompt for JSON, parse, validate, retry."""
    last_error = None
    for _ in range(attempts):
        resp = client.chat.completions.create(
            model="deeprelay/llama-3.3-70b-instruct",
            temperature=0,
            messages=[
                {"role": "system", "content": SYSTEM},
                {"role": "user", "content": question},
            ],
        )
        text = resp.choices[0].message.content.strip()
        # Models sometimes wrap JSON in a fence despite the instruction.
        if text.startswith("```"):
            text = text.strip("`").removeprefix("json").strip()
        try:
            data = json.loads(text)
            validate(instance=data, schema=SCHEMA)
            return data
        except (json.JSONDecodeError, ValidationError) as err:
            last_error = err
    raise RuntimeError(f"no valid JSON after {attempts} attempts: {last_error}")


print(structured("What is the capital of France?"))

Pick any served chat model from GET /v1/models?modality=chat — deeprelay/llama-3.3-70b-instruct above is one of them. temperature: 0 and a short, explicit schema in the system prompt materially cut the retry rate; jsonschema can be swapped for pydantic if you already model your responses there.

Roadmap

tools and tool_choice have shipped, streaming delta.tool_calls included — see Tool calling above.

response_format has not. It stays rejected rather than silently ignored, so you never get an unconstrained answer back from a request that asked for a constrained one. If it blocks you, tell us at support@deeprelay.ai — demand ordering is how this gets scheduled.

Errors

Errors use the OpenAI envelope, so SDK error handling works as-is (the /v1/messages endpoints use the Anthropic envelope instead; see Anthropic Messages API):

{ "error": { "message": "…", "type": "invalid_request_error", "code": "model_not_found" } }

Error responses also carry an X-Deeprelay-Error-Code header set to the same code (for example model_not_found, rate_limit_exceeded, insufficient_balance). The header uses one set of codes across every endpoint, so you can branch on or count errors without parsing the body. (An error that arrives mid-stream as an SSE error event carries its code in the event body only.)

Model availability. The serverless catalog is live, so a model id that worked earlier can stop being served. When that happens you get a model_not_found (404) with a "Try another model" message — re-fetch GET /v1/models for the current list and pick another. On the streaming endpoint this arrives as the final SSE error event before [DONE].

Concurrent-stream cap. Each API key may have a limited number of inference requests open at once. If you exceed it you get a 429 with code stream_limit_exceeded:

{ "error": { "message": "Too many concurrent streams open for this API key — close existing streams or retry shortly.", "type": "rate_limit_error", "code": "stream_limit_exceeded" } }

Close any streams you are no longer reading, or retry shortly — the slot frees as soon as an in-flight request finishes. This is separate from the per-second rate limit (rate_limit_exceeded), which bounds request rate rather than the number of simultaneously-open requests.

Upstream temporarily unavailable. If a model's serving capacity is degraded, requests may return a 503 with code upstream_unavailable:

{ "error": { "message": "Upstream service temporarily unavailable.", "type": "api_error", "code": "upstream_unavailable" } }

This is returned promptly (no long wait) while capacity recovers — retry after a short interval. The platform automatically routes around a degraded endpoint and resumes serving it once it recovers, so a 503 here is transient.

← All docs