model apis · chat / llm
Run Llama 3.3 70B Instruct by Meta through one OpenAI-compatible endpoint. Pay per use against your deeprelay balance. No instance to rent, no deployment to manage.
Available nowLive catalog rates, billed per use against your deeprelay balance, no minimums. Llama 3.3 70B Instruct is priced per use rather than included in the $2.99-a-month flat plan.
| Tier | Rate |
|---|---|
| Serverless | Input tokens: $1.10 / 1MOutput tokens: $1.10 / 1M |
rates are read every few minutes from the live model catalog and metered per request
Any OpenAI SDK works: set the base URL to https://api.deeprelay.ai/v1 and use your deeprelay API key.
curl https://api.deeprelay.ai/v1/chat/completions \
-H "Authorization: Bearer $DEEPRELAY_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "deeprelay/llama-3.3-70b-instruct",
"messages": [{"role": "user", "content": "Hello!"}]
}'from openai import OpenAI
client = OpenAI(
base_url="https://api.deeprelay.ai/v1",
api_key="YOUR_DEEPRELAY_API_KEY",
)
response = client.chat.completions.create(
model="deeprelay/llama-3.3-70b-instruct",
messages=[{"role": "user", "content": "Hello!"}],
)
print(response.choices[0].message.content)import OpenAI from "openai";
const client = new OpenAI({
baseURL: "https://api.deeprelay.ai/v1",
apiKey: process.env.DEEPRELAY_API_KEY,
});
const response = await client.chat.completions.create({
model: "deeprelay/llama-3.3-70b-instruct",
messages: [{ role: "user", content: "Hello!" }],
});
console.log(response.choices[0].message.content);| Model id | deeprelay/llama-3.3-70b-instruct |
| Modality | Chat / LLM |
| Author | Meta |
| Context window | 32,768 tokens |
| Streaming | Supported |
| Aliases | llama-3.3-70b, llama-3.3-70b-instruct, llama3.3-70b |
| Flat plan | Not included: pay as you go |
faq
Llama 3.3 70B Instruct on deeprelay currently costs $1.10 / 1M (input tokens), $1.10 / 1M (output tokens), billed against your deeprelay balance as you go. There are no minimums. It is priced per use rather than included in the $2.99-a-month flat plan, which covers a named set of models.
Yes. Point any OpenAI SDK at https://api.deeprelay.ai/v1 with a deeprelay API key and request model "deeprelay/llama-3.3-70b-instruct". No proprietary client needed.
Llama 3.3 70B Instruct supports a 32,768-token context window on deeprelay.
Chat models bill per input and output token at the listed per-million-token rates, metered per request against your deeprelay balance. Where a cached-input rate is listed, the part of a prompt the serving partner answers from its prompt cache (reported as prompt_tokens_details.cached_tokens) bills at that lower rate instead of the input rate.