300 million tokens. $2.99 a month.See the subscription  

model apis · chat / llm

GLM 5.2 API

Run GLM 5.2 by Z.ai through one OpenAI-compatible endpoint. Pay per use against your deeprelay balance. No instance to rent, no deployment to manage.

Available now

GLM 5.2 API pricing

Live catalog rates, billed per use against your deeprelay balance, no minimums. GLM 5.2 is priced per use rather than included in the $2.99-a-month flat plan.

TierRate
Serverless
Input tokens: $1.48 / 1MCached input tokens: $0.276 / 1MOutput tokens: $4.66 / 1M

rates are read every few minutes from the live model catalog and metered per request

Call GLM 5.2 in one request

Any OpenAI SDK works: set the base URL to https://api.deeprelay.ai/v1 and use your deeprelay API key.

curl
curl https://api.deeprelay.ai/v1/chat/completions \
  -H "Authorization: Bearer $DEEPRELAY_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "deeprelay/glm-5.2",
    "messages": [{"role": "user", "content": "Hello!"}]
  }'
Python (OpenAI SDK)
from openai import OpenAI

client = OpenAI(
    base_url="https://api.deeprelay.ai/v1",
    api_key="YOUR_DEEPRELAY_API_KEY",
)

response = client.chat.completions.create(
    model="deeprelay/glm-5.2",
    messages=[{"role": "user", "content": "Hello!"}],
)
print(response.choices[0].message.content)
JavaScript (OpenAI SDK)
import OpenAI from "openai";

const client = new OpenAI({
  baseURL: "https://api.deeprelay.ai/v1",
  apiKey: process.env.DEEPRELAY_API_KEY,
});

const response = await client.chat.completions.create({
  model: "deeprelay/glm-5.2",
  messages: [{ role: "user", content: "Hello!" }],
});
console.log(response.choices[0].message.content);

GLM 5.2 on deeprelay

Model iddeeprelay/glm-5.2
ModalityChat / LLM
AuthorZ.ai
Context window202,752 tokens
StreamingSupported
Aliasesglm-5.2, glm-5p2
Flat planNot included: pay as you go

faq

GLM 5.2 API questions

How much does the GLM 5.2 API cost?

GLM 5.2 on deeprelay currently costs $1.48 / 1M (input tokens), $0.276 / 1M (cached input tokens), $4.66 / 1M (output tokens), billed against your deeprelay balance as you go. There are no minimums. It is priced per use rather than included in the $2.99-a-month flat plan, which covers a named set of models.

Is the GLM 5.2 API OpenAI-compatible?

Yes. Point any OpenAI SDK at https://api.deeprelay.ai/v1 with a deeprelay API key and request model "deeprelay/glm-5.2". No proprietary client needed.

What is the context window of GLM 5.2?

GLM 5.2 supports a 202,752-token context window on deeprelay.

How does billing work?

Chat models bill per input and output token at the listed per-million-token rates, metered per request against your deeprelay balance. Where a cached-input rate is listed, the part of a prompt the serving partner answers from its prompt cache (reported as prompt_tokens_details.cached_tokens) bills at that lower rate instead of the input rate.