serverless inference
Every model behind our OpenAI-compatible inference API, priced per use. Point your existing SDK at api.deeprelay.ai/v1 and pay per token or per image. No instance to rent, no cold infrastructure to babysit.
| Model | Author | Context | Input / 1M | Cached input / 1M | Output / 1M |
|---|---|---|---|---|---|
| GLM 5.2 | Z.ai | 198K | $1.48 | $0.276 | $4.66 |
| GLM 5.3 | Z.ai | 198K | $1.48 | $0.276 | $4.66 |
| GLM 5.3 Flash | Z.ai | 1024K | $0.16 | $0.032 | $0.53 |
| GPT-OSS 120B | OpenAI | 128K | $0.16 | – | $0.64 |
| Kimi K2.6 | Moonshot AI | 256K | $1.01 | $0.17 | $4.24 |
| Kimi K2.7 Code | Moonshot AI | 256K | $1.01 | $0.201 | $4.24 |
| Kimi K3 | Moonshot AI | 1024K | $3.18 | $0.318 | $15.90 |
| Llama 3.3 70B Instruct | Meta | 32K | $1.10 | – | $1.10 |
| Minimax M3 | MiniMax | 512K | $0.32 | $0.064 | $1.27 |
| Muse Glimmer 30B | Meta | 128K | $0.37 | $0.042 | $1.59 |
| Qwen3.5 9B | Qwen | 256K | $0.18 | – | $0.27 |
| Qwen3.8 2.4t A95B | Qwen | 986K | $2.12 | $0.265 | $6.36 |
| Model | Author | Price |
|---|---|---|
| Flash Image 2.5 | $0.041 per image | |
| Flash Image 3.1 | $0.049 per image | |
| Flash Image 3.1 Lite | $0.073 per image | |
| FLUX.1 Kontext Max | Black Forest Labs | $0.085 per output megapixel |
| FLUX.1 Kontext Pro | Black Forest Labs | $0.042 per output megapixel |
| FLUX.1.1 Pro | Black Forest Labs | $0.042 per output megapixel |
| FLUX.2 Dev | Black Forest Labs | $0.016 per image |
| FLUX.2 Flex | Black Forest Labs | $0.032 per image |
| FLUX.2 Max | Black Forest Labs | $0.074 per output megapixel |
| FLUX.2 Pro | Black Forest Labs | $0.032 per image |
| Gemini 3 Pro Image | $0.142 per image | |
| GPT Image 1.5 | OpenAI | $0.036 per image |
| GPT Image 2 | OpenAI | $0.056 per image |
| Imagen 4.0 Fast | $0.021 per image | |
| Imagen 4.0 Preview | $0.042 per image | |
| Imagen 4.0 Ultra | $0.064 per image | |
| Qwen Image | Qwen | $0.0061 per image |
| Qwen Image 2.0 | Qwen | $0.037 per image |
| Qwen Image 2.0 Pro | Qwen | $0.08 per image |
| Seedream 3.0 | ByteDance | $0.019 per image |
| Seedream 4.0 | ByteDance | $0.032 per image |
| Seedream 5.0 Lite | ByteDance | $0.037 per image |
| Wan2.6 Image | Wan | $0.032 per image |
| Model | Author | Price |
|---|---|---|
| FLUX 3 | Black Forest Labs | $0.18 per second of video |
| Hailuo 02 | MiniMax | $0.059 per second of video |
| Happyhorse 1.0 T2v | Alibaba | $0.254 per second of video |
| Happyhorse 1.1 T2v | Alibaba | $0.148 per second of video |
| Kling 2.1 Standard | Kling | $0.039 per second of video |
| Minimax H3 | MiniMax | $0.147 per second of video |
| Pixverse V5 | PixVerse | $0.063 per second of video |
| Seedance 1.0 Lite | ByteDance | $0.03 per second of video |
| Seedance 1.0 Pro | ByteDance | $0.12 per second of video |
| Seedance 2.0 | ByteDance | $0.17 per second of video |
| Seedance 2.5 | ByteDance | $0.122 per second of video |
| Sora 2 | OpenAI | $0.106 per second of video |
| Sora 2 Pro | OpenAI | $0.398 per second of video |
| Veo 2.0 | $0.53 per second of video | |
| Vidu Q1 | Vidu | $0.047 per second of video |
| Vidu Q3 | Vidu | $0.103 per second of video |
| Vidu Q3 Turbo | Vidu | $0.207 per second of video |
| Wan2.7 T2v | Wan | $0.021 per second of video |
Every model is served through one OpenAI-compatible endpoint at https://api.deeprelay.ai/v1. Chat models bill per input and output token; the part of a prompt a partner serves from its prompt cache bills at the lower cached-input rate where one is listed (a dash means the model has no cached rate and every prompt token bills at the input rate). Image and video models bill per generation. Prices are the live catalog rates, the same figures the API and the console bill against: no minimums, billed against your deeprelay balance as you go. Models marked included ride the flat plan instead: $2.99 a month carries 300M input and 30M output tokens across those models, weighted by each model's usage factor, and anything past the allowance falls back to the rates above. Models marked economy trade a cold start of roughly 30–60 seconds for a lower rate. This page refreshes every few minutes from the live model catalog.