Model overview
Doubling is the whole story here: $1.90 per million input tokens instead of $0.95, $8 output instead of $4, $0.38 on cache hits instead of $0.19. Nothing else moves. The window stays at 262,144 tokens, and reasoning, tools, multimodal input and context caching are unchanged from the standard Code listing. What the surcharge buys is serving throughput, which makes this a latency question rather than a capability one. It pays for itself in interactive editors, chat surfaces and agent loops where wall-clock time carries its own cost, and argues against itself in batch jobs or overnight evaluation runs where a queue absorbs the delay. Budget twice, then measure whether the faster response changes anything downstream.
Billing conditions: Standard Kimi API price. Cache-miss input is shown as Input; automatic cache-hit input is shown as Cached input.
- Context window
- 262,144 tokens
- Knowledge cutoff
- —
- Price last verified
- First observed
Descriptions, capabilities, and tariffs come from the linked source. First-party prices take priority; OpenRouter fallbacks are labelled.