An xAI estimate needs the model, prompt-length pricing tier and any tool or media operations. Use the live rows below to separate those charges and compare editable workload assumptions.

How xAI pricing works

Applicable text models have short- and long-context pricing. When a prompt reaches the documented long-context boundary, the corresponding rates apply to the request’s tokens, not only the portion beyond the boundary. Cached input remains a distinct charge category. Official documentation.

Check the tier against the entire assembled prompt. Conversation history, document excerpts and tool definitions can change its length even when the latest user message is short. Compare alternatives under the same request shape so one model is not given an artificially smaller workload.

Server-side tools add operation charges alongside token usage. A tool-using request can perform more than one operation before returning, so its final answer length alone does not explain the bill. Official documentation.

xAI pricing structure: context-conditioned input, cache, output and separate tool/media charges.

Full price list

Verified model prices
ModelPrice typeUSDUnitTierRegionSource
grok-4.20-0309-non-reasoningCached Input$0.4per 1M tokenslong contextSee sourceOfficial source ↗
grok-4.20-0309-non-reasoningCached Input$0.4per 1M tokenslong context ≥200kSee sourceOfficial source ↗
grok-4.20-0309-non-reasoningCached Input$0.2per 1M tokensshort contextSee sourceOfficial source ↗
grok-4.20-0309-non-reasoningCached Input$0.2per 1M tokensshort context <200kSee sourceOfficial source ↗
grok-4.20-0309-non-reasoningInput$2.5per 1M tokenslong contextSee sourceOfficial source ↗
grok-4.20-0309-non-reasoningInput$2.5per 1M tokenslong context ≥200kSee sourceOfficial source ↗
grok-4.20-0309-non-reasoningInput$1.25per 1M tokensshort contextSee sourceOfficial source ↗
grok-4.20-0309-non-reasoningInput$1.25per 1M tokensshort context <200kSee sourceOfficial source ↗
grok-4.20-0309-non-reasoningOutput$5per 1M tokenslong contextSee sourceOfficial source ↗
grok-4.20-0309-non-reasoningOutput$5per 1M tokenslong context ≥200kSee sourceOfficial source ↗
grok-4.20-0309-non-reasoningOutput$2.5per 1M tokensshort contextSee sourceOfficial source ↗
grok-4.20-0309-non-reasoningOutput$2.5per 1M tokensshort context <200kSee sourceOfficial source ↗
grok-4.20-0309-reasoningCached Input$0.4per 1M tokenslong contextSee sourceOfficial source ↗
grok-4.20-0309-reasoningCached Input$0.4per 1M tokenslong context ≥200kSee sourceOfficial source ↗
grok-4.20-0309-reasoningCached Input$0.2per 1M tokensshort contextSee sourceOfficial source ↗
grok-4.20-0309-reasoningCached Input$0.2per 1M tokensshort context <200kSee sourceOfficial source ↗
grok-4.20-0309-reasoningInput$2.5per 1M tokenslong contextSee sourceOfficial source ↗
grok-4.20-0309-reasoningInput$2.5per 1M tokenslong context ≥200kSee sourceOfficial source ↗
grok-4.20-0309-reasoningInput$1.25per 1M tokensshort contextSee sourceOfficial source ↗
grok-4.20-0309-reasoningInput$1.25per 1M tokensshort context <200kSee sourceOfficial source ↗
grok-4.20-0309-reasoningOutput$5per 1M tokenslong contextSee sourceOfficial source ↗
grok-4.20-0309-reasoningOutput$5per 1M tokenslong context ≥200kSee sourceOfficial source ↗
grok-4.20-0309-reasoningOutput$2.5per 1M tokensshort contextSee sourceOfficial source ↗
grok-4.20-0309-reasoningOutput$2.5per 1M tokensshort context <200kSee sourceOfficial source ↗
grok-4.20-multi-agent-0309Cached Input$0.4per 1M tokenslong contextSee sourceOfficial source ↗
grok-4.20-multi-agent-0309Cached Input$0.4per 1M tokenslong context ≥200kSee sourceOfficial source ↗
grok-4.20-multi-agent-0309Cached Input$0.2per 1M tokensshort contextSee sourceOfficial source ↗
grok-4.20-multi-agent-0309Cached Input$0.2per 1M tokensshort context <200kSee sourceOfficial source ↗
grok-4.20-multi-agent-0309Input$2.5per 1M tokenslong contextSee sourceOfficial source ↗
grok-4.20-multi-agent-0309Input$2.5per 1M tokenslong context ≥200kSee sourceOfficial source ↗
grok-4.20-multi-agent-0309Input$1.25per 1M tokensshort contextSee sourceOfficial source ↗
grok-4.20-multi-agent-0309Input$1.25per 1M tokensshort context <200kSee sourceOfficial source ↗
grok-4.20-multi-agent-0309Output$5per 1M tokenslong contextSee sourceOfficial source ↗
grok-4.20-multi-agent-0309Output$5per 1M tokenslong context ≥200kSee sourceOfficial source ↗
grok-4.20-multi-agent-0309Output$2.5per 1M tokensshort contextSee sourceOfficial source ↗
grok-4.20-multi-agent-0309Output$2.5per 1M tokensshort context <200kSee sourceOfficial source ↗
grok-4.3Cached Input$0.4per 1M tokenslong contextSee sourceOfficial source ↗
grok-4.3Cached Input$0.4per 1M tokenslong context ≥200kSee sourceOfficial source ↗
grok-4.3Cached Input$0.2per 1M tokensshort contextSee sourceOfficial source ↗
grok-4.3Cached Input$0.2per 1M tokensshort context <200kSee sourceOfficial source ↗
grok-4.3Input$2.5per 1M tokenslong contextSee sourceOfficial source ↗
grok-4.3Input$2.5per 1M tokenslong context ≥200kSee sourceOfficial source ↗
grok-4.3Input$1.25per 1M tokensshort contextSee sourceOfficial source ↗
grok-4.3Input$1.25per 1M tokensshort context <200kSee sourceOfficial source ↗
grok-4.3Output$5per 1M tokenslong contextSee sourceOfficial source ↗
grok-4.3Output$5per 1M tokenslong context ≥200kSee sourceOfficial source ↗
grok-4.3Output$2.5per 1M tokensshort contextSee sourceOfficial source ↗
grok-4.3Output$2.5per 1M tokensshort context <200kSee sourceOfficial source ↗
grok-4.5Cached Input$0.6per 1M tokenslong contextSee sourceOfficial source ↗
grok-4.5Cached Input$0.6per 1M tokenslong context ≥200kSee sourceOfficial source ↗
grok-4.5Cached Input$0.3per 1M tokensshort contextSee sourceOfficial source ↗
grok-4.5Cached Input$0.3per 1M tokensshort context <200kSee sourceOfficial source ↗
grok-4.5Input$4per 1M tokenslong contextSee sourceOfficial source ↗
grok-4.5Input$4per 1M tokenslong context ≥200kSee sourceOfficial source ↗
grok-4.5Input$2per 1M tokensshort contextSee sourceOfficial source ↗
grok-4.5Input$2per 1M tokensshort context <200kSee sourceOfficial source ↗
grok-4.5Output$12per 1M tokenslong contextSee sourceOfficial source ↗
grok-4.5Output$12per 1M tokenslong context ≥200kSee sourceOfficial source ↗
grok-4.5Output$6per 1M tokensshort contextSee sourceOfficial source ↗
grok-4.5Output$6per 1M tokensshort context <200kSee sourceOfficial source ↗
grok-4.6Cached Input$1per 1M tokenslong contextSee sourceOfficial source ↗
grok-4.6Cached Input$1per 1M tokenslong context ≥200kSee sourceOfficial source ↗
grok-4.6Cached Input$0.5per 1M tokensshort contextSee sourceOfficial source ↗
grok-4.6Cached Input$0.5per 1M tokensshort context <200kSee sourceOfficial source ↗
grok-4.6Input$4per 1M tokenslong contextSee sourceOfficial source ↗
grok-4.6Input$4per 1M tokenslong context ≥200kSee sourceOfficial source ↗
grok-4.6Input$2per 1M tokensshort contextSee sourceOfficial source ↗
grok-4.6Input$2per 1M tokensshort context <200kSee sourceOfficial source ↗
grok-4.6Output$12per 1M tokenslong contextSee sourceOfficial source ↗
grok-4.6Output$12per 1M tokenslong context ≥200kSee sourceOfficial source ↗
grok-4.6Output$6per 1M tokensshort contextSee sourceOfficial source ↗
grok-4.6Output$6per 1M tokensshort context <200kSee sourceOfficial source ↗
grok-build-0.1Cached Input$0.4per 1M tokenslong contextSee sourceOfficial source ↗
grok-build-0.1Cached Input$0.4per 1M tokenslong context ≥200kSee sourceOfficial source ↗
grok-build-0.1Cached Input$0.2per 1M tokensshort contextSee sourceOfficial source ↗
grok-build-0.1Cached Input$0.2per 1M tokensshort context <200kSee sourceOfficial source ↗
grok-build-0.1Input$2per 1M tokenslong contextSee sourceOfficial source ↗
grok-build-0.1Input$2per 1M tokenslong context ≥200kSee sourceOfficial source ↗
grok-build-0.1Input$1per 1M tokensshort contextSee sourceOfficial source ↗
grok-build-0.1Input$1per 1M tokensshort context <200kSee sourceOfficial source ↗
grok-build-0.1Output$4per 1M tokenslong contextSee sourceOfficial source ↗
grok-build-0.1Output$4per 1M tokenslong context ≥200kSee sourceOfficial source ↗
grok-build-0.1Output$2per 1M tokensshort contextSee sourceOfficial source ↗
grok-build-0.1Output$2per 1M tokensshort context <200kSee sourceOfficial source ↗

Last verified · Source ↗

The rows preserve billing types and conditions. Select the tier your request qualifies for rather than treating the lowest input row as a universal price. Missing values require source verification; they are not zero-cost operations.

For image, video or audio work, read the media unit explicitly. A duration-based rate and a token rate cannot be ranked as though they described the same quantity. Keep the requested output properties with your estimate.

Estimate your cost

Interactive tool

Estimate your API costs

Your text and estimates stay in this browser. No API requests are sent to model providers.

Loading verified model records…

Enter observed prompt and output usage from a representative call. The calculator’s defaults are illustrative. It estimates supported token charges, so list excluded tool and media operations separately before approving a budget.

Three worked scenarios

Chatbot. Retained messages can make later turns more expensive than the initial greeting. Compare your actual history policy and observe cache behavior across turns.

Interactive tool

Estimate your API costs

Your text and estimates stay in this browser. No API requests are sent to model providers.

Loading verified model records…

Summarization. Prompt assembly can cross a pricing boundary. Compare a focused document excerpt with the complete source before choosing the context tier.

Interactive tool

Estimate your API costs

Your text and estimates stay in this browser. No API requests are sent to model providers.

Loading verified model records…

Coding assistant. Include repository context, diagnostic logs and follow-up reasoning. If tools are enabled, add their observed operations to the token-only estimate.

Interactive tool

Estimate your API costs

Your text and estimates stay in this browser. No API requests are sent to model providers.

Loading verified model records…

How to reduce cost on xAI

xAI caches matching starting messages. Its guidance recommends a stable conversation identifier and unchanged earlier messages; changing or reordering a prefix can remove the expected benefit. Cache reuse should be measured, not assumed. Official documentation.

Cached usage appears in different fields for Chat Completions and Responses. Read the field appropriate to your endpoint when estimating the cached share. Official documentation.

The Batch API supports asynchronous work on eligible models. Review each model’s support and the batch request contract before assigning an offline evaluation job to that path. Official documentation.

Control tool scope as carefully as prompt length. Ask whether a request needs current retrieval at all, then inspect which tools were actually invoked. Evaluate a cheaper candidate on the full task, including failed attempts, rather than reduce cost by removing information required for a correct answer.

Billing, credits and limits notes

Team prepaid credits and monthly invoiced billing are separate payment arrangements. Auto top-up can buy additional credits, while the invoiced spending control governs a different part of the payment flow. Inspect both before assuming a balance is a fixed budget. Official documentation.

Match the console’s selected team to the key used by the application. Review the xAI free-access evidence and team rate limits when a request cannot proceed. Adding credit does not repair an invalid payload.

Price change history

  1. xAI · grok-build-0.1 — MetadataRecord updated; consult the linked source for details. → Record updated; consult the linked source for details.Source ↗
  2. xAI · grok-4.6 — MetadataRecord updated; consult the linked source for details. → Record updated; consult the linked source for details.Source ↗
  3. xAI · grok-4.5 — MetadataRecord updated; consult the linked source for details. → Record updated; consult the linked source for details.Source ↗
  4. xAI · grok-4.3 — MetadataRecord updated; consult the linked source for details. → Record updated; consult the linked source for details.Source ↗
  5. xAI · grok-4.20-multi-agent-0309 — MetadataRecord updated; consult the linked source for details. → Record updated; consult the linked source for details.Source ↗
  6. xAI · grok-4.20-0309-reasoning — MetadataRecord updated; consult the linked source for details. → Record updated; consult the linked source for details.Source ↗
  7. xAI · grok-4.20-0309-non-reasoning — MetadataRecord updated; consult the linked source for details. → Record updated; consult the linked source for details.Source ↗
  8. xAI · grok-build-0.1 — MetadataRecord updated; consult the linked source for details. → Record updated; consult the linked source for details.Source ↗
  9. xAI · grok-4.6 — MetadataRecord updated; consult the linked source for details. → Record updated; consult the linked source for details.Source ↗
  10. xAI · grok-4.5 — MetadataRecord updated; consult the linked source for details. → Record updated; consult the linked source for details.Source ↗

Subscribe to the changelog RSS feed

Our history begins with recorded observations. It does not reconstruct all earlier prices. For billing reconciliation, retain actual per-request usage and the provider’s invoice or credit records with the model and task identifiers.

Use the AI API cost calculator to turn the model and workload you are considering into an estimate.

Last verified · Source ↗

Frequently asked questions

Are long-context rates charged only on excess tokens?
For applicable models, xAI states that the long-context rates apply to the request’s tokens once its prompt reaches the boundary. Official documentation.
Can tool usage make a short answer expensive?
Yes. A request can include tool operations and internal token work before its final answer. Inspect actual request cost. Official documentation.
Where is cached usage in Responses?
The documentation places it within input_tokens_details.cached_tokens. Chat Completions uses its prompt-token details object. Official documentation.
Is the credit balance a complete spending cap?
Check auto top-up and invoiced billing as well as the prepaid balance. These controls govern different payment paths. Official documentation.
What should I compare when changing models?
Use identical representative tasks, record completion quality and include retries or corrections. Compare the cost of useful work rather than only listed input rates.

Sources

Last verified · Source ↗