Model overview
Paying three times more per token buys no extra context or capability on paper: deepseek-v4-pro publishes the same 1M-token window and the same reasoning, tool-calling, structured-output and fill-in-the-middle support as the cheaper tier, at $1.32 per million input tokens and $3.96 output. One anchor helps size the gap — $1.32 is also what Flash charges to generate a million tokens, so reading a long document here costs what writing it there would. Cached input falls to $0.044, thirty times under the miss rate, so prefix-heavy agent loops absorb the premium better than one-shot calls. These are peak rates; every one halves outside the 01:00-04:00 and 06:00-10:00 UTC windows.
Billing conditions: Standard peak price. Cache-miss input is shown as Input; cache-hit input is shown as Cached input. DeepSeek halves every rate off-peak.
- Context window
- 1,000,000 tokens
- Knowledge cutoff
- —
- Price last verified
- First observed
Descriptions, capabilities, and tariffs come from the linked source. First-party prices take priority; OpenRouter fallbacks are labelled.