Build a Kimi API budget from uncached input, reused context and generated output. Keep optional tools and the actual processing mode attached to the estimate.
How Moonshot pricing works
Kimi bills model input and output; document content becomes model input when it is passed into an inference request. Official documentation.
Describe each stage of the application before counting. Uploading a file, extracting its content, asking a question and revising an answer are different actions. The useful budget follows the actions the application actually takes. If a pipeline repeats the same document across several questions, preserve that repetition in the workload and investigate whether the cache evidence supports a different input category.
Kimi K3 billing is described as flat across context length, with separate cache-hit input, cache-miss input and output categories. Official documentation.
Use the live table for amounts and applicable conditions. A flat context schedule does not mean that adding irrelevant material is free: the amount of billable input still matters. Likewise, a low input estimate can hide an expensive response policy. Define the required answer before forecasting its length, then inspect accepted results rather than assuming every response reaches the configured cap.

Full price list
| Model | Price type | USD | Unit | Tier | Region | Source |
|---|---|---|---|---|---|---|
| kimi-k2.6 | Cached Input | $0.16 | per 1M tokens | Standard | See source | Official source ↗ |
| kimi-k2.6 | Input | $0.95 | per 1M tokens | Standard | See source | Official source ↗ |
| kimi-k2.6 | Output | $4 | per 1M tokens | Standard | See source | Official source ↗ |
| kimi-k2.7-code | Cached Input | $0.19 | per 1M tokens | Standard | See source | Official source ↗ |
| kimi-k2.7-code | Input | $0.95 | per 1M tokens | Standard | See source | Official source ↗ |
| kimi-k2.7-code | Output | $4 | per 1M tokens | Standard | See source | Official source ↗ |
| kimi-k2.7-code-highspeed | Cached Input | $0.38 | per 1M tokens | Standard | See source | Official source ↗ |
| kimi-k2.7-code-highspeed | Input | $1.9 | per 1M tokens | Standard | See source | Official source ↗ |
| kimi-k2.7-code-highspeed | Output | $8 | per 1M tokens | Standard | See source | Official source ↗ |
| kimi-k3 | Cached Input | $0.3 | per 1M tokens | Standard | See source | Official source ↗ |
| kimi-k3 | Input | $3 | per 1M tokens | Standard | See source | Official source ↗ |
| kimi-k3 | Output | $15 | per 1M tokens | Standard | See source | Official source ↗ |
Last verified · Source ↗
Read the model identity and charge category together. If a required hosted feature is not represented, keep it outside the token estimate as an unresolved category instead of silently assigning it no cost.
Estimate your cost
Estimate your API costs
Your text and estimates stay in this browser. No API requests are sent to model providers.
Loading verified model records…
Save a Kimi workload in the API cost calculator and preserve its cache and output assumptions.
Compare an uncached baseline with a measured-reuse scenario. A difference between those estimates is useful only when the request structure makes the reuse plausible. If the application frequently changes its initial context, use the conservative case until the returned usage supports a better assumption.
Three worked scenarios
For a chatbot, keep the stable instructions separate from the changing conversation. Inspect whether the application repeatedly sends history that no longer helps answer the current question. Reduce that material only after checking the effect on accepted responses. The preset below supplies a starting shape; replace its assumptions with a safe representative conversation.
Estimate your API costs
Your text and estimates stay in this browser. No API requests are sent to model providers.
Loading verified model records…
For summarization, measure the document material and the final summary independently. If a first pass creates section summaries and a second pass combines them, include both. Keep a fixture with important facts near the end of the document so cost reduction does not quietly remove necessary evidence.
Estimate your API costs
Your text and estimates stay in this browser. No API requests are sent to model providers.
Loading verified model records…
For a coding assistant, include the project context, generated patch and any additional model attempts needed after checks fail. Keep failed attempts visible in cost per completed change. A cheaper individual call is not necessarily a cheaper accepted patch when the workflow requires repeated repair.
Estimate your API costs
Your text and estimates stay in this browser. No API requests are sent to model providers.
Loading verified model records…
How to reduce cost on Moonshot
Kimi caching is automatic for eligible repeated initial context; the application does not create a cache identifier or manage its lifetime. Official documentation.
Keep truly stable material consistent and place changing task input after it where the request design allows. Avoid cosmetic prompt churn during the measurement because it makes the reason for a changed cache result harder to isolate. Confirm the answer still uses the current source material when a document changes; preserving an obsolete prompt to chase reuse would defeat the application’s purpose.
The Batch API provides a JSONL submission and results workflow for asynchronous inference. Official documentation.
Use that route when delayed completion fits the job. Preserve a custom identifier for each input and reconcile every result, including failures. Test a small batch before submitting the full workload so a formatting error does not become a large unresolved queue. Compare the final accepted outputs with the synchronous baseline before changing the production schedule.
Billing, credits and limits notes
Keep recharge history, remaining balance, promotional conditions and the application forecast as separate records. A voucher does not demonstrate an ongoing free entitlement. Before unattended work, decide what happens when the project’s configured budget rejects requests: pause the queue, retain its unfinished items and notify the responsible operator through your application’s normal process.
Review Moonshot offer conditions and capacity planning with the same model and project. Do not solve a capacity diagnosis by adding funds until the error evidence identifies the account condition.
Price change history
- Moonshot — ValueNot previously recorded → Tier: Tier5 · Metric: TPM · Value: 5000000 · Notes: Published account tier. Unlimited quotas are not converted into a numeric cap.Source ↗
- Moonshot — ValueNot previously recorded → Tier: Tier5 · Metric: RPM · Value: 300 · Notes: Published account tier. Unlimited quotas are not converted into a numeric cap.Source ↗
- Moonshot — ValueNot previously recorded → Tier: Tier5 · Metric: concurrency · Value: 100 · Notes: Published account tier. Unlimited quotas are not converted into a numeric cap.Source ↗
- Moonshot — ValueNot previously recorded → Tier: Tier4 · Metric: TPM · Value: 4000000 · Notes: Published account tier. Unlimited quotas are not converted into a numeric cap.Source ↗
- Moonshot — ValueNot previously recorded → Tier: Tier4 · Metric: RPM · Value: 200 · Notes: Published account tier. Unlimited quotas are not converted into a numeric cap.Source ↗
- Moonshot — ValueNot previously recorded → Tier: Tier4 · Metric: concurrency · Value: 60 · Notes: Published account tier. Unlimited quotas are not converted into a numeric cap.Source ↗
- Moonshot — ValueNot previously recorded → Tier: Tier3 · Metric: TPM · Value: 3000000 · Notes: Published account tier. Unlimited quotas are not converted into a numeric cap.Source ↗
- Moonshot — ValueNot previously recorded → Tier: Tier3 · Metric: RPM · Value: 200 · Notes: Published account tier. Unlimited quotas are not converted into a numeric cap.Source ↗
- Moonshot — ValueNot previously recorded → Tier: Tier3 · Metric: concurrency · Value: 50 · Notes: Published account tier. Unlimited quotas are not converted into a numeric cap.Source ↗
- Moonshot — ValueNot previously recorded → Tier: Tier2 · Metric: TPM · Value: 3000000 · Notes: Published account tier. Unlimited quotas are not converted into a numeric cap.Source ↗
When a recorded rate changes, rerun the previous assumptions before also changing the prompt. This separates the provider-rate effect from the application-workload effect. Keep the old decision record so the change can be explained.
Last verified · Source ↗
Frequently asked questions
Is extracted document text part of inference input?
Does K3 use a context-length pricing tier?
Do I create a cache resource for ordinary Kimi prefix reuse?
Can I use a batch estimate for an interactive call?
What should a coding budget include?
Should an unknown tool charge be treated as free?
Sources
- Kimi API quickstart ↗
- Kimi model list ↗
- Inference pricing ↗
- Recharge and limits ↗
- Error reference ↗
- Organization management ↗
- Account and billing ↗
- Context caching ↗
- Batch API ↗
- Model parameters ↗
Last verified · Source ↗