Plan Groq traffic around organization capacity, project controls and the specific model being called. This guide explains how to identify the exhausted allowance without confusing a short burst with a billing problem.

How Groq limits work

Groq applies organization-level limits and allows more restrictive project settings. A request can encounter a request-count or token constraint first, and some accounts have separate input and output token controls. Official documentation.

Measure the work queued by the whole application, including retries and background tasks. A per-user throttle is insufficient if several users share one organization. Keep an application-wide scheduler that can pause or reduce concurrency when the shared capacity becomes scarce.

Project customization does not create capacity outside the organization’s ceiling. Set a development project’s limits deliberately so a load test cannot silently consume all the shared allocation. Official documentation.

Groq rate-limit scopes: organization ceiling, project restrictions and model-specific request/token checks.

Limits by tier and model

Verified rate limits
Model or scopeTierMetricLimitNotesSource
canopylabs/orpheus-arabic-saudi
canopylabs/orpheus-arabic-saudi
freeRPD100Published plan limit; account overrides may apply.Official source ↗
canopylabs/orpheus-arabic-saudi
canopylabs/orpheus-arabic-saudi
freeRPM10Published plan limit; account overrides may apply.Official source ↗
canopylabs/orpheus-arabic-saudi
canopylabs/orpheus-arabic-saudi
freeTPD3,600Published plan limit; account overrides may apply.Official source ↗
canopylabs/orpheus-arabic-saudi
canopylabs/orpheus-arabic-saudi
freeTPM1,200Published plan limit; account overrides may apply.Official source ↗
canopylabs/orpheus-v1-english
canopylabs/orpheus-v1-english
freeRPD100Published plan limit; account overrides may apply.Official source ↗
canopylabs/orpheus-v1-english
canopylabs/orpheus-v1-english
freeRPM10Published plan limit; account overrides may apply.Official source ↗
canopylabs/orpheus-v1-english
canopylabs/orpheus-v1-english
freeTPD3,600Published plan limit; account overrides may apply.Official source ↗
canopylabs/orpheus-v1-english
canopylabs/orpheus-v1-english
freeTPM1,200Published plan limit; account overrides may apply.Official source ↗
GPT OSS 120B
openai/gpt-oss-120b
freeRPD1,000Published plan limit; account overrides may apply.Official source ↗
GPT OSS 120B
openai/gpt-oss-120b
freeRPM30Published plan limit; account overrides may apply.Official source ↗
GPT OSS 120B
openai/gpt-oss-120b
freeTPD200,000Published plan limit; account overrides may apply.Official source ↗
GPT OSS 120B
openai/gpt-oss-120b
freeTPM8,000Published plan limit; account overrides may apply.Official source ↗
GPT OSS 20B
openai/gpt-oss-20b
freeRPD1,000Published plan limit; account overrides may apply.Official source ↗
GPT OSS 20B
openai/gpt-oss-20b
freeRPM30Published plan limit; account overrides may apply.Official source ↗
GPT OSS 20B
openai/gpt-oss-20b
freeTPD200,000Published plan limit; account overrides may apply.Official source ↗
GPT OSS 20B
openai/gpt-oss-20b
freeTPM8,000Published plan limit; account overrides may apply.Official source ↗
groq/compound
groq/compound
freeRPD250Published plan limit; account overrides may apply.Official source ↗
groq/compound
groq/compound
freeRPM30Published plan limit; account overrides may apply.Official source ↗
groq/compound
groq/compound
freeTPM70,000Published plan limit; account overrides may apply.Official source ↗
groq/compound-mini
groq/compound-mini
freeRPD250Published plan limit; account overrides may apply.Official source ↗
groq/compound-mini
groq/compound-mini
freeRPM30Published plan limit; account overrides may apply.Official source ↗
groq/compound-mini
groq/compound-mini
freeTPM70,000Published plan limit; account overrides may apply.Official source ↗
Llama Prompt Guard 2 22M
meta-llama/llama-prompt-guard-2-22m
freeRPD14,400Published plan limit; account overrides may apply.Official source ↗
Llama Prompt Guard 2 22M
meta-llama/llama-prompt-guard-2-22m
freeRPM30Published plan limit; account overrides may apply.Official source ↗
Llama Prompt Guard 2 22M
meta-llama/llama-prompt-guard-2-22m
freeTPD500,000Published plan limit; account overrides may apply.Official source ↗
Llama Prompt Guard 2 22M
meta-llama/llama-prompt-guard-2-22m
freeTPM15,000Published plan limit; account overrides may apply.Official source ↗
Prompt Guard 2 86M
meta-llama/llama-prompt-guard-2-86m
freeRPD14,400Published plan limit; account overrides may apply.Official source ↗
Prompt Guard 2 86M
meta-llama/llama-prompt-guard-2-86m
freeRPM30Published plan limit; account overrides may apply.Official source ↗
Prompt Guard 2 86M
meta-llama/llama-prompt-guard-2-86m
freeTPD500,000Published plan limit; account overrides may apply.Official source ↗
Prompt Guard 2 86M
meta-llama/llama-prompt-guard-2-86m
freeTPM15,000Published plan limit; account overrides may apply.Official source ↗
Qwen/Qwen3.6-27B
qwen/qwen3.6-27b
freeRPD1,000Published plan limit; account overrides may apply.Official source ↗
Qwen/Qwen3.6-27B
qwen/qwen3.6-27b
freeRPM30Published plan limit; account overrides may apply.Official source ↗
Qwen/Qwen3.6-27B
qwen/qwen3.6-27b
freeTPD200,000Published plan limit; account overrides may apply.Official source ↗
Qwen/Qwen3.6-27B
qwen/qwen3.6-27b
freeTPM8,000Published plan limit; account overrides may apply.Official source ↗
Qwen/Qwen3.8-27B
qwen/qwen3.8-27b
freeRPD1,000Published plan limit; account overrides may apply.Official source ↗
Qwen/Qwen3.8-27B
qwen/qwen3.8-27b
freeRPM30Published plan limit; account overrides may apply.Official source ↗
Qwen/Qwen3.8-27B
qwen/qwen3.8-27b
freeTPD200,000Published plan limit; account overrides may apply.Official source ↗
Qwen/Qwen3.8-27B
qwen/qwen3.8-27b
freeTPM8,000Published plan limit; account overrides may apply.Official source ↗
Safety GPT OSS 20B
openai/gpt-oss-safeguard-20b
freeRPD1,000Published plan limit; account overrides may apply.Official source ↗
Safety GPT OSS 20B
openai/gpt-oss-safeguard-20b
freeRPM30Published plan limit; account overrides may apply.Official source ↗
Safety GPT OSS 20B
openai/gpt-oss-safeguard-20b
freeTPD200,000Published plan limit; account overrides may apply.Official source ↗
Safety GPT OSS 20B
openai/gpt-oss-safeguard-20b
freeTPM8,000Published plan limit; account overrides may apply.Official source ↗
Whisper
whisper-large-v3
freeRPD2,000Published plan limit; account overrides may apply.Official source ↗
Whisper
whisper-large-v3
freeRPM20Published plan limit; account overrides may apply.Official source ↗
Whisper Large V3 Turbo
whisper-large-v3-turbo
freeRPD2,000Published plan limit; account overrides may apply.Official source ↗
Whisper Large V3 Turbo
whisper-large-v3-turbo
freeRPM20Published plan limit; account overrides may apply.Official source ↗

Last verified · Source ↗

Use the live table to find the relevant model and plan, then compare it with your organization’s console. Missing values should be checked at the source before a launch decision; they are not unlimited capacity. Distinguish a token metric from an audio-duration metric.

For capacity planning, retain request timestamps, model identifiers and observed usage. Aggregate them over the provider’s actual windows. This makes it possible to explain whether long prompts, many short requests or overlapping workers caused the rejection.

How to move up a tier

The Developer plan provides access to higher capacity and features such as Batch and Flex. Groq’s documented plan limits are a baseline; specialized workloads can require a separate capacity discussion. Official documentation.

Explain the workload when requesting capacity: interactive or offline, typical prompt shape, output length, peak concurrency and acceptable waiting time. A precise request gives the provider a better basis for guidance than asking for unlimited access.

Check Groq paid usage and Free-plan constraints before upgrading. More capacity can let the same worker spend its budget faster, so update the cost guard at the same time.

Reading limit headers and 429 responses

Groq’s request headers describe daily request allowance, while token headers describe per-minute token allowance. The reset fields use their respective windows. The retry-after header is supplied on a rate-limited response. Official documentation.

  • x-ratelimit-remaining-requests: inspect the request allowance, not an assumed minute-only counter.
  • x-ratelimit-remaining-tokens: inspect remaining token capacity for the documented window.
  • x-ratelimit-reset-requests and x-ratelimit-reset-tokens: preserve both when diagnosing a burst.
  • retry-after: use the supplied delay when handling a rate rejection.

Record headers beside the rejected request’s workload characteristics. Do not overwrite the first failure with only the final retry result; that loses the evidence explaining why the queue stalled.

Retry strategy the provider recommends

Groq’s error guidance calls for throttling and respecting rate limits. Its Python SDK has automatic retries for selected failures, so inspect that behavior before adding another retry layer. Official documentation.

Use bounded attempts, respect the server’s timing and introduce jitter so several workers do not resume together. Keep the overall task deadline explicit. A user-facing action should either finish, report that it is still queued or fail with a recoverable state; it should not remain inside hidden retry loops indefinitely.

Do not retry a permissions error as if it were capacity. Preserve the failure class, pause the affected work and resolve the organization or project control. Repeatedly sending an unchanged forbidden request cannot fix its policy.

Batch/async options that bypass limits

Groq provides Batch for asynchronous workloads. This is a different processing path with its own documented eligibility and constraints, not an exemption from all limits. Use it for work whose results can be collected later. Official documentation.

Prompt caching can change repeated-input accounting on supported models. Cache credit is applied after processing, so simultaneous large requests can still encounter capacity limits. Official documentation.

Separate the interactive and offline queues. Give each a visible completion state and a task identifier that survives retries. This prevents a deferred job from blocking an interactive user while making eventual reconciliation straightforward.

Use the AI API cost calculator to turn the model and workload you are considering into an estimate.

Last verified · Source ↗

Frequently asked questions

Do Groq request-limit headers always mean requests per minute?
No. The documentation identifies request headers with the daily request allowance and token headers with the token-per-minute allowance. Official documentation.
Can I exceed organization capacity by splitting work across projects?
No. Project limits are bounded by organization capacity. Official documentation.
Why can a cached workload still hit a limit?
Groq notes that cached tokens are subtracted after processing. Parallel requests can encounter capacity before that adjustment. Official documentation.
Does Batch mean unlimited work?
No. Batch is an asynchronous processing option with its own constraints. Review the supported request format and limits. Official documentation.
Should I add retries around SDK retries?
Only with a clear shared attempt budget and deadline. Stacked retries can make a small failure generate much more traffic than intended.

Sources

Last verified · Source ↗