Model overview
Several coordinated passes run behind one endpoint here, and the rate card gives no hint of it: $1.25 per million input tokens, $0.20 cached and $2.50 output over 1,000,000 tokens of context, matching both single-model 4.20 builds and grok-4-3 exactly. The multiplication happens in consumption, not in the unit price. Each internal pass re-reads its context and emits its own tokens, so a single request can bill a multiple of what the identically priced non-reasoning variant would charge for the same question. Two levers matter as a result: the $0.20 cached rate, a sixth of the miss price, which pays off when every pass replays the same prefix, and the 200,000-token threshold above which xAI's long-context rates apply to the whole request.
Billing conditions: Standard short-context API price. At 200K input tokens or more, xAI applies its published long-context rates to all tokens in the request.
- Context window
- 1,000,000 tokens
- Knowledge cutoff
- —
- Price last verified
- First observed
Descriptions, capabilities, and tariffs come from the linked source. First-party prices take priority; OpenRouter fallbacks are labelled.