Budget Model Studio requests by model, deployment scope and billable category. Keep context tiers, caching, batch processing and media units attached to the actual request path.
How Alibaba pricing works
The inference pricing reference separates models, deployment scopes, input conditions and charge units, including model-specific context tiers. Official documentation.
Begin with the operation the application sends. A text request may use input and output categories, while an image, speech or video feature can use a different unit. Keep that unit visible throughout the estimate. Do not convert a media price into an assumed text-token rate merely to make a table look uniform.
Select the applicable deployment scope before comparing candidates. A model name alone does not identify every condition behind a price row. Keep the actual workspace host and model with the workload record so an apparent discrepancy can be checked against the correct source.
For tiered context pricing, preserve the threshold condition in the data row rather than quoting an unconditional rate in prose. Estimate the assembled input, including instructions and retrieved material, then inspect whether the workload crosses a qualifying boundary. A small prompt demonstration may not represent a document-heavy application.

Full price list
| Model | Price type | USD | Unit | Tier | Region | Source |
|---|---|---|---|---|---|---|
| qwen3.8-max | Input | $2 | per 1M tokens | 0<Token≤1M | international | Official source ↗ |
| qwen3.8-max | Output | $6 | per 1M tokens | 0<Token≤1M | international | Official source ↗ |
Last verified · Source ↗
Use the source and condition attached to each row. If the application requires a category that is not recorded, retain it as an unresolved item. Missing information should never be interpreted as free usage.
Estimate your cost
Estimate your API costs
Your text and estimates stay in this browser. No API requests are sent to model providers.
Loading verified model records…
Use the Alibaba API cost calculator for supported token categories and keep other units in a separate estimate.
Create an ordinary and a demanding scenario using the same accepted-output requirement. If a candidate needs another inference stage to repair or validate its answer, include that stage. This gives a more useful comparison than the price of an isolated call that does not complete the task.
Three worked scenarios
For a chatbot, include the conversation material that is actually resent and the length of an accepted answer. Check whether unnecessary history can be removed without breaking continuity. Keep prompt revisions and cache assumptions separate so their effect can be measured.
Estimate your API costs
Your text and estimates stay in this browser. No API requests are sent to model providers.
Loading verified model records…
For summarization, preserve the source document, extraction or split stages and the final synthesis. If the application needs field-level evidence, count the structure required to return it. A cheaper short summary is not an equivalent result when it omits the facts the workflow must retain.
Estimate your API costs
Your text and estimates stay in this browser. No API requests are sent to model providers.
Loading verified model records…
For a coding assistant, estimate supplied project context and generated changes, then include failed attempts that required another model call. Validate a patch with the project’s checks before treating the task as accepted. Keep the chosen coding or general-purpose model explicit in the scenario.
Estimate your API costs
Your text and estimates stay in this browser. No API requests are sent to model providers.
Loading verified model records…
How to reduce cost on Alibaba
Model Studio documents explicit and implicit context caching as different mechanisms with model-specific support and mutually exclusive use. Official documentation.
Choose the mechanism that fits the request path and inspect the returned usage. A request meeting a technical eligibility condition does not prove it produced a cache hit. Keep a cold baseline and a measured reuse case, and preserve the prefix while investigating changes. Do not retain stale documents solely to increase reuse.
The compatible Batch API supplies an asynchronous workflow for supported inference requests. Official documentation.
Use batch only when delayed results fit the application. Start with a small file, assign stable custom request identities and reconcile every output and error row. Compare accepted results with the synchronous baseline before changing the workload schedule.
For a simpler task, test a less costly model with the same output contract. Keep the difficult cases in the evaluation. If a cheaper candidate fails missing-evidence or structured-output cases, either restrict it to the tasks it passes or retain the stronger candidate for that workflow.
Billing, credits and limits notes
The billing guide provides usage and cost-management views, while the free-quota rules exclude several non-real-time operations from the offer. Official documentation.
Reconcile the provider bill against the application stages rather than comparing it with a single token total. Keep optional cloud storage, deployment or training activities outside the inference estimate unless they are explicitly part of the plan. A model’s free quota does not establish that every associated cloud operation is covered.
Review free-quota scope and account throughput before running a recurring job. Decide how work stops at the intended spending boundary.
Price change history
- Alibaba Cloud — TermsNot previously recorded → Terms: Eligible newly activated Model Studio accounts can receive model-specific token quotas. Eligibility, model coverage, and expiration are defined by the official quota page. · Expiry days: 90Source ↗
- Alibaba Cloud · qwen-plus-character-ja — ValueNot previously recorded → Tier: international default · Metric: TPM · Value: 500000 · Notes: International deployment scope. Models with dynamic quotas require an account check.Source ↗
- Alibaba Cloud · qwen-plus-character-ja — ValueNot previously recorded → Tier: international default · Metric: RPM · Value: 120 · Notes: International deployment scope. Models with dynamic quotas require an account check.Source ↗
- Alibaba Cloud · qwen-flash-character — ValueNot previously recorded → Tier: international default · Metric: TPM · Value: 500000 · Notes: International deployment scope. Models with dynamic quotas require an account check.Source ↗
- Alibaba Cloud · qwen-flash-character — ValueNot previously recorded → Tier: international default · Metric: RPM · Value: 120 · Notes: International deployment scope. Models with dynamic quotas require an account check.Source ↗
- Alibaba Cloud · qwen-plus-character — ValueNot previously recorded → Tier: international default · Metric: TPM · Value: 500000 · Notes: International deployment scope. Models with dynamic quotas require an account check.Source ↗
- Alibaba Cloud · qwen-plus-character — ValueNot previously recorded → Tier: international default · Metric: RPM · Value: 120 · Notes: International deployment scope. Models with dynamic quotas require an account check.Source ↗
- Alibaba Cloud · text-embedding-v3 — ValueNot previously recorded → Tier: international default · Metric: TPM · Value: 24000000 · Notes: International deployment scope. Models with dynamic quotas require an account check.Source ↗
- Alibaba Cloud · text-embedding-v3 — ValueNot previously recorded → Tier: international default · Metric: RPM · Value: 6000 · Notes: International deployment scope. Models with dynamic quotas require an account check.Source ↗
- Alibaba Cloud · text-embedding-v4 — ValueNot previously recorded → Tier: international default · Metric: TPM · Value: 1000000 · Notes: International deployment scope. Models with dynamic quotas require an account check.Source ↗
When a price or tier condition changes, apply it to the previous workload assumptions first. Then evaluate prompt, model or processing-mode changes separately so the final estimate remains explainable.
Last verified · Source ↗
Frequently asked questions
Does a model name identify the complete price condition?
Can I compare speech or image work as ordinary text tokens?
Does cache eligibility guarantee a hit?
Is batch a discount for the same interactive request?
Does free quota cover every associated cloud operation?
What makes a cheaper model an acceptable replacement?
Sources
- First Qwen API call ↗
- Recommended models ↗
- Model inference pricing ↗
- API keys and permissions ↗
- Rate limits ↗
- Error codes ↗
- New-user free quota ↗
- Model usage ↗
- Context cache ↗
- Batch API ↗
- Billing and cost management ↗
- Dynamic rate limiting ↗
- Qwen Coder ↗
- Model updates ↗
Last verified · Source ↗