Model overview
deepseek-v4-flash is the cheaper of DeepSeek's two published tiers: $0.44 per million input tokens and $1.32 output, exactly one third of what the Pro tier charges on both sides. The two models declare the same 1M-token context window and the same feature set — reasoning, tool calls, structured output and fill-in-the-middle — so the choice between them is a price decision rather than a capability one. Cache hits are billed at $0.014, roughly a thirtieth of the cache-miss rate, which matters for repeated prompts sharing a long prefix. The rates above are peak: DeepSeek halves every one of them outside the 01:00-04:00 and 06:00-10:00 UTC windows.
Billing conditions: Standard peak price. Cache-miss input is shown as Input; cache-hit input is shown as Cached input. DeepSeek halves every rate off-peak.
- Context window
- 1,000,000 tokens
- Knowledge cutoff
- —
- Price last verified
- First observed
Descriptions, capabilities, and tariffs come from the linked source. First-party prices take priority; OpenRouter fallbacks are labelled.