Budget Model Studio requests by model, deployment scope and billable category. Keep context tiers, caching, batch processing and media units attached to the actual request path.

How Alibaba pricing works

The inference pricing reference separates models, deployment scopes, input conditions and charge units, including model-specific context tiers. Official documentation.

Begin with the operation the application sends. A text request may use input and output categories, while an image, speech or video feature can use a different unit. Keep that unit visible throughout the estimate. Do not convert a media price into an assumed text-token rate merely to make a table look uniform.

Select the applicable deployment scope before comparing candidates. A model name alone does not identify every condition behind a price row. Keep the actual workspace host and model with the workload record so an apparent discrepancy can be checked against the correct source.

For tiered context pricing, preserve the threshold condition in the data row rather than quoting an unconditional rate in prose. Estimate the assembled input, including instructions and retrieved material, then inspect whether the workload crosses a qualifying boundary. A small prompt demonstration may not represent a document-heavy application.

PNG separating model/scope, context tier, input/output, cache mechanism, batch and media units without hardcoded amounts.

Full price list

Verified model prices
ModelPrice typeUSDUnitTierRegionSource
qwen3.8-maxInput$2per 1M tokens0<Token≤1MinternationalOfficial source ↗
qwen3.8-maxOutput$6per 1M tokens0<Token≤1MinternationalOfficial source ↗

Last verified · Source ↗

Use the source and condition attached to each row. If the application requires a category that is not recorded, retain it as an unresolved item. Missing information should never be interpreted as free usage.

Estimate your cost

Interactive tool

Estimate your API costs

Your text and estimates stay in this browser. No API requests are sent to model providers.

Loading verified model records…

Use the Alibaba API cost calculator for supported token categories and keep other units in a separate estimate.

Create an ordinary and a demanding scenario using the same accepted-output requirement. If a candidate needs another inference stage to repair or validate its answer, include that stage. This gives a more useful comparison than the price of an isolated call that does not complete the task.

Three worked scenarios

For a chatbot, include the conversation material that is actually resent and the length of an accepted answer. Check whether unnecessary history can be removed without breaking continuity. Keep prompt revisions and cache assumptions separate so their effect can be measured.

Interactive tool

Estimate your API costs

Your text and estimates stay in this browser. No API requests are sent to model providers.

Loading verified model records…

For summarization, preserve the source document, extraction or split stages and the final synthesis. If the application needs field-level evidence, count the structure required to return it. A cheaper short summary is not an equivalent result when it omits the facts the workflow must retain.

Interactive tool

Estimate your API costs

Your text and estimates stay in this browser. No API requests are sent to model providers.

Loading verified model records…

For a coding assistant, estimate supplied project context and generated changes, then include failed attempts that required another model call. Validate a patch with the project’s checks before treating the task as accepted. Keep the chosen coding or general-purpose model explicit in the scenario.

Interactive tool

Estimate your API costs

Your text and estimates stay in this browser. No API requests are sent to model providers.

Loading verified model records…

How to reduce cost on Alibaba

Model Studio documents explicit and implicit context caching as different mechanisms with model-specific support and mutually exclusive use. Official documentation.

Choose the mechanism that fits the request path and inspect the returned usage. A request meeting a technical eligibility condition does not prove it produced a cache hit. Keep a cold baseline and a measured reuse case, and preserve the prefix while investigating changes. Do not retain stale documents solely to increase reuse.

The compatible Batch API supplies an asynchronous workflow for supported inference requests. Official documentation.

Use batch only when delayed results fit the application. Start with a small file, assign stable custom request identities and reconcile every output and error row. Compare accepted results with the synchronous baseline before changing the workload schedule.

For a simpler task, test a less costly model with the same output contract. Keep the difficult cases in the evaluation. If a cheaper candidate fails missing-evidence or structured-output cases, either restrict it to the tasks it passes or retain the stronger candidate for that workflow.

Billing, credits and limits notes

The billing guide provides usage and cost-management views, while the free-quota rules exclude several non-real-time operations from the offer. Official documentation.

Reconcile the provider bill against the application stages rather than comparing it with a single token total. Keep optional cloud storage, deployment or training activities outside the inference estimate unless they are explicitly part of the plan. A model’s free quota does not establish that every associated cloud operation is covered.

Review free-quota scope and account throughput before running a recurring job. Decide how work stops at the intended spending boundary.

Price change history

  1. Alibaba Cloud — TermsNot previously recorded → Terms: Eligible newly activated Model Studio accounts can receive model-specific token quotas. Eligibility, model coverage, and expiration are defined by the official quota page. · Expiry days: 90Source ↗
  2. Alibaba Cloud · qwen-plus-character-ja — ValueNot previously recorded → Tier: international default · Metric: TPM · Value: 500000 · Notes: International deployment scope. Models with dynamic quotas require an account check.Source ↗
  3. Alibaba Cloud · qwen-plus-character-ja — ValueNot previously recorded → Tier: international default · Metric: RPM · Value: 120 · Notes: International deployment scope. Models with dynamic quotas require an account check.Source ↗
  4. Alibaba Cloud · qwen-flash-character — ValueNot previously recorded → Tier: international default · Metric: TPM · Value: 500000 · Notes: International deployment scope. Models with dynamic quotas require an account check.Source ↗
  5. Alibaba Cloud · qwen-flash-character — ValueNot previously recorded → Tier: international default · Metric: RPM · Value: 120 · Notes: International deployment scope. Models with dynamic quotas require an account check.Source ↗
  6. Alibaba Cloud · qwen-plus-character — ValueNot previously recorded → Tier: international default · Metric: TPM · Value: 500000 · Notes: International deployment scope. Models with dynamic quotas require an account check.Source ↗
  7. Alibaba Cloud · qwen-plus-character — ValueNot previously recorded → Tier: international default · Metric: RPM · Value: 120 · Notes: International deployment scope. Models with dynamic quotas require an account check.Source ↗
  8. Alibaba Cloud · text-embedding-v3 — ValueNot previously recorded → Tier: international default · Metric: TPM · Value: 24000000 · Notes: International deployment scope. Models with dynamic quotas require an account check.Source ↗
  9. Alibaba Cloud · text-embedding-v3 — ValueNot previously recorded → Tier: international default · Metric: RPM · Value: 6000 · Notes: International deployment scope. Models with dynamic quotas require an account check.Source ↗
  10. Alibaba Cloud · text-embedding-v4 — ValueNot previously recorded → Tier: international default · Metric: TPM · Value: 1000000 · Notes: International deployment scope. Models with dynamic quotas require an account check.Source ↗

Subscribe to the changelog RSS feed

When a price or tier condition changes, apply it to the previous workload assumptions first. Then evaluate prompt, model or processing-mode changes separately so the final estimate remains explainable.

Last verified · Source ↗

Frequently asked questions

Does a model name identify the complete price condition?
No. Preserve the deployment scope, unit and any qualifying tier.
Can I compare speech or image work as ordinary text tokens?
Only if the official accounting explicitly supports that conversion; otherwise keep the original unit.
Does cache eligibility guarantee a hit?
No. Inspect actual returned usage and keep an uncached baseline.
Is batch a discount for the same interactive request?
It is a distinct asynchronous workflow that must fit the application’s result timing.
Does free quota cover every associated cloud operation?
No. Inspect the specific offer scope and keep excluded activities separate.
What makes a cheaper model an acceptable replacement?
It must pass the same task-specific output checks and operating requirements.

Sources

Last verified · Source ↗