Model overview
Direct answers with no intermediate chain are the point of this build, and they come at no discount: $1.25 in, $0.20 cached, $2.50 out and a 1,000,000-token window, the same four numbers the reasoning and multi-agent builds carry. Since xAI charges by token rather than by mode, the saving is real but indirect, fewer generated tokens at the same $2.50 rate on tasks where deliberation adds nothing. That makes it the sensible default for extraction, classification, formatting and routing, with the reasoning build reserved for problems that actually fail without it. Note that the published capability list is unchanged as well: reasoning parameters, tool calls and structured output are declared here too.
Billing conditions: Standard short-context API price. At 200K input tokens or more, xAI applies its published long-context rates to all tokens in the request.
- Context window
- 1,000,000 tokens
- Knowledge cutoff
- —
- Price last verified
- First observed
Descriptions, capabilities, and tariffs come from the linked source. First-party prices take priority; OpenRouter fallbacks are labelled.