Model overview
Deliberation before answering is the only thing that sets this build apart, because the meter does not move: $1.25 per million input tokens, $0.20 cached and $2.50 output are identical to its non-reasoning twin, to the multi-agent build and to grok-4-3, all on the same 1,000,000-token window. The cost difference therefore lives entirely in token volume. Reasoning traces are billed at the output rate, twice the input rate, so a short answer can cost several times its visible length once the chain is counted. Pick it when correctness on multi-step problems matters more than a predictable invoice, and measure actual output counts rather than the rate card, which will show four identical rows.
Billing conditions: Standard short-context API price. At 200K input tokens or more, xAI applies its published long-context rates to all tokens in the request.
- Context window
- 1,000,000 tokens
- Knowledge cutoff
- —
- Price last verified
- First observed
Descriptions, capabilities, and tariffs come from the linked source. First-party prices take priority; OpenRouter fallbacks are labelled.