Use a consistent price baseline to find low-cost text-model candidates. The cheapest application depends on its accepted outputs, request shape and extra work required to complete each task.

What matters for cheapest ai apis

Begin with the output contract. A chatbot answer, document extract and code patch require different generated work, so one price ordering cannot represent every application equally. Decide what counts as an accepted result before testing a less costly model.

The baseline ranking uses the specification’s input-heavy blend to make text candidates comparable. It is a reference assumption, not a measurement of your application. A response-heavy workflow can order the same candidates differently. Keep the blend visible in the rendered data and replace it with a measured scenario for the final decision.

Provider pricing separates input and output categories, with additional conditions for supported modes and features. Official documentation.

Inspect missing categories and qualifying tiers. A zero amount supported by a source is different from an unknown price. Exclude incompatible units from a token comparison and keep optional hosted-feature charges outside a partial token estimate until they are established.

Ranked candidates

Only active candidates with documented values for this ranking are included. Prices retain the tier and deployment condition shown below.

Blended price = (input price × 3 + output price) ÷ 4, in USD per 1M tokens. This fixed mix is a comparison measure; estimate your own workload separately.

Models ranked by blended price
ModelProviderInput USD / 1MOutput USD / 1MBlended USD / 1MContext tokensPrice conditionCost
Cohere: North Mini Code (free)OpenRouterFreeFreeFree256,000StandardEstimate cost
Dots Studio: Dots3-Note Preview (free)OpenRouterFreeFreeFree512,000StandardEstimate cost
Free Models RouterOpenRouterFreeFreeFree200,000StandardEstimate cost
Gemini 2.5 Flash Preview TTSGoogleFreeFreeFree8,192free tierEstimate cost
Gemini 3.1 Flash TTS PreviewGoogleFreeFreeFree8,192free tierEstimate cost
Gemini 3.5 Live TranslateGoogleFreeFreeFree16,384free tierEstimate cost
Gemini 3.5 TranscribeGoogleFreeFreeFree98,304free tierEstimate cost
Gemini 3.5 Transcribe LiveGoogleFreeFreeFree131,072free tierEstimate cost
Google: Gemma 4 26B A4B (free)OpenRouterFreeFreeFree262,144StandardEstimate cost
Google: Gemma 4 31B (free)OpenRouterFreeFreeFree262,144StandardEstimate cost
Google: Lyria 3 Clip PreviewOpenRouterFreeFreeFree1,048,576StandardEstimate cost
Google: Lyria 3 Pro PreviewOpenRouterFreeFreeFree1,048,576StandardEstimate cost
LiquidAI: LFM2.5-2.6B (free)OpenRouterFreeFreeFree65,536StandardEstimate cost
NVIDIA: Nemotron 3 Nano Omni (free)OpenRouterFreeFreeFree256,000StandardEstimate cost
NVIDIA: Nemotron 3 Super (free)OpenRouterFreeFreeFree262,144StandardEstimate cost
NVIDIA: Nemotron 3 Ultra (free)OpenRouterFreeFreeFree1,000,000StandardEstimate cost
NVIDIA: Nemotron 3.5 Content Safety (free)OpenRouterFreeFreeFree128,000StandardEstimate cost
NVIDIA: Nemotron 3.5 Lightning (free)OpenRouterFreeFreeFree1,000,000StandardEstimate cost
Nex AGI: Nex-N2.5-Mini (free)OpenRouterFreeFreeFree262,144StandardEstimate cost
Nex AGI: Nex-N2.5-Pro (free)OpenRouterFreeFreeFree262,144StandardEstimate cost

Last verified · Source ↗

Read the source dates and full provider pricing before deciding. The table narrows the candidate set; it does not certify account access, output quality or the cost of a multi-stage application. Preserve the host and model identifier with every copied row.

Our three picks

Interactive tool

Find your starting point

Your text and estimates stay in this browser. No API requests are sent to model providers.

Loading verified model records…

Use the cheapest role to choose the first low-rate candidate to test. Use balanced to investigate whether a different documented fit reduces repair work. Use strongest only with suitable comparative evidence, or treat it as an explicit evaluation candidate when that evidence is missing.

Keep the safe task fixture unchanged across these roles. If a candidate needs another call to repair invalid output, record the additional work. Do not make the cheaper candidate appear successful by weakening the response requirement or removing difficult cases from its test.

When each pick is wrong

The lowest rate is a poor choice when the model cannot meet a mandatory input or output requirement. It can also be a poor choice when response repair, human checking or failed downstream processing erases the apparent saving. Measure those consequences on representative work.

A balanced or quality-oriented choice is unnecessary for a narrow task that the cheaper candidate reliably passes. Avoid paying for a broader capability merely because its description sounds more advanced. Route simple tasks deliberately and keep a regression fixture that verifies the boundary.

For a specialized task, compare coding candidates, chatbot requirements or embedding models before applying a general text-price ordering.

Estimate cost

Interactive tool

Estimate your API costs

Your text and estimates stay in this browser. No API requests are sent to model providers.

Loading verified model records…

Open the cost calculator and enter your actual input/output mix, request volume and supported cache assumptions.

Compare ordinary and demanding tasks. Retain the same acceptance standard while changing the candidate. Review the provider’s full categories for any tool or media feature outside the table, then reconcile the forecast against returned usage after deployment.

Last verified · Source ↗

Frequently asked questions

Does the cheapest token rate mean lowest application cost?
No. Extra calls, repair work and other charge categories can change the result.
Why use a blended baseline?
It provides a consistent starting comparison; your actual input/output mix can differ.
Is an unknown price free?
No. Missing evidence remains unresolved.
How should I compare a smaller model?
Use the same fixture and acceptance criteria, including difficult cases.
What belongs in cost per accepted task?
All model stages and repair work required to complete the task successfully.
When should I prefer the cheaper candidate?
When it passes the required task and operating checks at lower complete workload cost.

Sources

Last verified · Source ↗