Pricing
A flat monthly subscription over the models we serve direct, with pay-as-you-go for the rest of the catalog. No seats, no minimums, and no daily or hourly caps. Every number here is one you can watch move in the console.
One flat price for the models we serve direct, sized for real production traffic rather than a trial.
Models included
Billed monthly, renews until cancelled. Cancel any time and access runs to the end of the paid period. See the Terms of Service.
More tiers
Larger allowances and team plans are on the way.
How it is counted
The allowance is 300 million input and 30 million output tokens a month, with no daily or hourly caps. Usage breathes inside a 75 million-token week. Heavier models draw it faster through a published usage factor, and limit changes are posted before they apply. The live meter sits in your console.
Pay as you go
The live catalog rates, the same figures the API and the console bill against. Chat and embedding models bill per token, image models per image or per output megapixel, video models per second of output. No minimums and no reservation: add credit, send requests, pay for what the request used.
| Model | Context | Input / 1M | Cached input / 1M | Output / 1M |
|---|---|---|---|---|
| GLM 5.2Z.ai · deeprelay/glm-5.2 | 198K | $1.48 | $0.276 | $4.66 |
| GLM 5.3Z.ai · deeprelay/glm-5.3 | 198K | $1.48 | $0.276 | $4.66 |
| GLM 5.3 FlashZ.ai · deeprelay/glm-5.3-flash | 1M | $0.16 | $0.032 | $0.53 |
| GPT-OSS 120BOpenAI · deeprelay/gpt-oss-120b | 128K | $0.16 | – | $0.64 |
| Kimi K2.6Moonshot AI · deeprelay/kimi-k2.6 | 256K | $1.01 | $0.17 | $4.24 |
| Kimi K2.7 CodeMoonshot AI · deeprelay/kimi-k2.7-code | 256K | $1.01 | $0.201 | $4.24 |
| Kimi K3Moonshot AI · deeprelay/kimi-k3 | 1M | $3.18 | $0.318 | $15.90 |
| Llama 3.3 70B InstructMeta · deeprelay/llama-3.3-70b-instruct | 32K | $1.10 | – | $1.10 |
| Minimax M3MiniMax · deeprelay/minimax-m3 | 512K | $0.32 | $0.064 | $1.27 |
| Muse Glimmer 30BMeta · deeprelay/muse-glimmer-30b | 128K | $0.37 | $0.042 | $1.59 |
| Qwen3.5 9BQwen · deeprelay/qwen3.5-9b | 256K | $0.18 | – | $0.27 |
| Qwen3.8 2.4t A95BQwen · deeprelay/qwen3.8-2.4t-a95b | 1M | $2.12 | $0.265 | $6.36 |
| Model | Context | Input / 1M |
|---|---|---|
| Qwen3 Embedding 8BQwen · deeprelay/qwen3-embedding-8b | 40K | $0.11 |
Prices are USD per 1M tokens for chat and embeddings, and per unit of output for media. The cached input column is the rate for the part of a prompt served from a partner's prompt cache, typically far below the input rate; a dash there means the model publishes no cached rate, so every prompt token bills at the input rate. A dash anywhere is an absent rate, never a zero. An economy row is the same model on cheaper capacity, trading a cold start of roughly 30 to 60 seconds for the lower rate. This page refreshes from the catalog every few minutes.
Questions
Requests keep working. Past the allowance they bill pay-as-you-go at the published per-model rate, the same rates listed below. Nothing stops and nothing queues.
No, and the rate tables below say which do. The plan covers the models we serve direct; every other model in the catalog bills pay-as-you-go at its listed rate.
Each model carries a published usage factor. A model at ×2 draws two tokens of allowance per token served. The factors ship with the plan terms and changes are posted before they apply, so there are no silent multipliers.
None. The only shaping is a 75 million-token weekly fair-use ceiling, so a burst on Tuesday doesn't cost you Wednesday.
Yes, from the billing portal in your console. Your plan stays active to the end of the period you already paid for, and renewal simply doesn't happen.
No. It is the same OpenAI-compatible endpoint either way: same paths, same request bodies, same streaming frames. The plan changes what you are billed, never how you call it.