300 million tokens. $2.99 a month.See the subscription  

Pricing

One plan. Every model on it.

A flat monthly subscription over the models we serve direct, with pay-as-you-go for the rest of the catalog. No seats, no minimums, and no daily or hourly caps. Every number here is one you can watch move in the console.

Base

Flat rate

One flat price for the models we serve direct, sized for real production traffic rather than a trial.

$2.99per month
  • 300M input tokens a month
  • 30M output tokens a month
  • No daily or hourly caps
  • 75M-token weekly fair-use ceiling

Models included

  • DeepSeek V4 Pro
  • DeepSeek V4 Flash
  • DeepSeek V4.1 Flash

Billed monthly, renews until cancelled. Cancel any time and access runs to the end of the paid period. See the Terms of Service.

More tiers

Larger allowances and team plans are on the way.

How it is counted

A mechanism, not a meter.

The allowance is 300 million input and 30 million output tokens a month, with no daily or hourly caps. Usage breathes inside a 75 million-token week. Heavier models draw it faster through a published usage factor, and limit changes are posted before they apply. The live meter sits in your console.

Pay as you go

Every model, at its rate.

The live catalog rates, the same figures the API and the console bill against. Chat and embedding models bill per token, image models per image or per output megapixel, video models per second of output. No minimums and no reservation: add credit, send requests, pay for what the request used.

Embeddings

1 model

Prices are USD per 1M tokens for chat and embeddings, and per unit of output for media. The cached input column is the rate for the part of a prompt served from a partner's prompt cache, typically far below the input rate; a dash there means the model publishes no cached rate, so every prompt token bills at the input rate. A dash anywhere is an absent rate, never a zero. An economy row is the same model on cheaper capacity, trading a cold start of roughly 30 to 60 seconds for the lower rate. This page refreshes from the catalog every few minutes.

Questions

What happens when I use up the monthly allowance?

Requests keep working. Past the allowance they bill pay-as-you-go at the published per-model rate, the same rates listed below. Nothing stops and nothing queues.

Do all models come with the plan?

No, and the rate tables below say which do. The plan covers the models we serve direct; every other model in the catalog bills pay-as-you-go at its listed rate.

Why do some models draw the allowance faster?

Each model carries a published usage factor. A model at ×2 draws two tokens of allowance per token served. The factors ship with the plan terms and changes are posted before they apply, so there are no silent multipliers.

Are there daily or hourly caps?

None. The only shaping is a 75 million-token weekly fair-use ceiling, so a burst on Tuesday doesn't cost you Wednesday.

Can I cancel any time?

Yes, from the billing portal in your console. Your plan stays active to the end of the period you already paid for, and renewal simply doesn't happen.

Is the API different on the plan?

No. It is the same OpenAI-compatible endpoint either way: same paths, same request bodies, same streaming frames. The plan changes what you are billed, never how you call it.