AI Tools
Groq pricing in 2026: what each model costs, and what moved to contact sales
Groq pricing in September 2026 starts at $0.075 per 1M input tokens for GPT OSS 20B and $0.15 for GPT OSS 120B, with output at $0.30 and $0.60. Llama 3.1 8B and Llama 3.3 70B no longer carry published prices: both rows now read Contact Sales.
What matters
- GPT OSS 120B costs $0.15 input and $0.60 output per 1M tokens on Groq, and GPT OSS 20B costs $0.075/$0.30, verified on the GroqCloud models page on September 21, 2026.
- Llama 3.1 8B, Llama 3.3 70B and MiniMax M2.7 sit in the same tables as Enterprise models with Contact Sales in both the price and rate-limit columns.
- Whisper is billed per audio hour, not per token: $0.111 for whisper-large-v3 and $0.04 for whisper-large-v3-turbo.
- Free-plan limits for the open models are 30 requests per minute, 1,000 requests per day, 8K tokens per minute and 200K tokens per day.
- Cerebras lists the same GPT OSS 120B at $0.35/$0.75 per 1M tokens at about 3,000 tokens per second, so Groq's edge is price per token, not raw speed.
Groq pricing per model, verified September 21, 2026
Groq's machine-readable rate card lives in the developer docs at console.groq.com/docs/models, not on the marketing site. That table is the single public source for what Groq charges, and it is what this article verified on September 21, 2026.
Two production models carry per-token prices. GPT OSS 120B (openai/gpt-oss-120b) is $0.15 input and $0.60 output per 1M tokens at 500 tokens per second. GPT OSS 20B (openai/gpt-oss-20b) is $0.075 and $0.30 at 1,000 tokens per second. Both list a 131,072-token context window and a 65,536-token maximum completion, with developer-plan limits of 250K tokens per minute and 1,000 requests per minute.
Speech is metered by time: whisper-large-v3 costs $0.111 per hour of audio and whisper-large-v3-turbo $0.04 per hour. In the preview section, Qwen3.8-27B (qwen/qwen3.8-27b) is the one text model with a published rate at $0.80 input and $4.00 output per 1M tokens, and the Orpheus voices bill per character at $22 per 1M for English and $40 per 1M for Arabic Saudi.
What the Contact Sales rows mean for buyers
Llama 3.1 8B (llama-3.1-8b-instant) and Llama 3.3 70B (llama-3.3-70b-versatile) appear in Groq's Production Models table with the Enterprise label, 560 and 280 tokens per second respectively, and Contact Sales where the price and rate limits should be. MiniMax M2.7 is listed the same way.
Groq's changelog describes how those enterprise models work: MiniMax M2.5 and Qwen3-VL 32B were announced as available for Enterprise customers, with access arranged through a Groq account team rather than self-serve checkout. A second page corroborates the Llama rows: the per-model table on Groq's rate-limits page lists no Llama model at all, free plan or otherwise.
This is a recent change, not a permanent feature of the docs. An Internet Archive capture of the same page from August 21, 2026 shows the Production Models table with no Llama rows at all, while the September 5, 2026 capture shows both rows present and already reading Contact Sales. The practical effect is that the shrinkage happened during August and early September 2026, after most tutorials and comparison posts were written.
Groq free tier vs developer plan limits
The rate-limits page opens on the free plan view. On the free plan, the GPT OSS models and Qwen3.8-27B cap at 30 requests per minute, 1,000 requests per day, 8K tokens per minute and 200K tokens per day. Whisper caps at 20 requests per minute, 2,000 requests per day, 7.2K audio seconds per hour and 28.8K audio seconds per day. The Orpheus voices are the tightest, at 10 requests per minute, 100 requests per day and 1.2K tokens per minute.
Paying lifts those ceilings. The models table quotes developer-plan limits per model: 250K tokens per minute and 1,000 requests per minute for GPT OSS 120B, GPT OSS 20B and Qwen3.8-27B, 150K tokens per minute for GPT OSS Safeguard 20B, and 50K tokens per minute for the Orpheus voices. Groq also states that cached tokens do not count against any limit, that limits apply at the organization level rather than per API key, and that exceeding a limit returns HTTP 429 with a retry-after header.
The number to read twice is 8K tokens per minute. Any request larger than that cannot clear the free plan's token budget at all, which makes the free tier a testing surface for small prompts rather than a production option for long-context work.
Groq vs Cerebras vs OpenRouter on the same models
Speed is what the per-token price buys. Cerebras lists openai/gpt-oss-120b at $0.35 input and $0.75 output per 1M tokens at roughly 3,000 tokens per second, so Groq's $0.15/$0.60 is less than half the input price for about one sixth of the throughput. On Qwen 3.8 27B the two split differently: Groq charges $0.80/$4.00 at 450 tokens per second, Cerebras $0.99/$1.49 at about 1,850 tokens per second.
OpenRouter's public model list is the third reference point, because it shows what the same model IDs trade at across providers. It lists openai/gpt-oss-120b at $0.15 input and $0.60 output per 1M tokens, identical to Groq's own list price, openai/gpt-oss-20b at $0.03/$0.13, below Groq's $0.075/$0.30, and qwen/qwen3.8-27b at $0.20/$2.55 with a 1M-token context window listed.
Put together, the three tables say Groq sells fast access to a small set of open models at commodity prices. It is not the cheapest per token in the market, and it is not the fastest, which is exactly why the price and the speed columns have to be read together.
Who should pay Groq prices, and who should not
- Good fit: interactive products that need sub-second replies on GPT OSS or Qwen models and stay inside a 131,072-token context with completions up to 65,536 tokens.
- Bad fit: anything pinned to Llama 3.1 8B, Llama 3.3 70B or MiniMax M2.7. Those now need an enterprise conversation, and OpenRouter still lists public prices for Llama 3.3 70B at $0.10/$0.32 per 1M tokens.
- Watch the meter if your load is bursty: on the free plan the 8K tokens-per-minute ceiling, not the dollar price, is what breaks first.
- Steady, batch-tolerant volume is a different problem. Self-hosting open weights on a budget VPS such as RackNerd's annual KVM plans or a Hostinger VPS removes the per-token meter entirely, and trades it for operations work you own.
Limitations that show up after you commit
- Qwen3.8-27B lives in the preview section, and Groq's own note says preview models may be discontinued at short notice.
- Qwen3.8-27B caps completions at 16,384 tokens and uploads at 20 MB, well below the 65,536-token completions on the GPT OSS models.
- Prompt Guard and the Orpheus voices bill by character or by hour, so they do not map onto a cost-per-1M-tokens calculation at all.
- Higher limits exist only for select workloads and enterprise use cases, which means the listed developer numbers are a floor for serious volume, not a ceiling.
At a glance
| Model | Price per 1M tokens | Speed | Developer-plan limits |
|---|---|---|---|
| GPT OSS 120B | $0.15 input / $0.60 output | 500 t/s | 250K TPM, 1K RPM |
| GPT OSS 20B | $0.075 input / $0.30 output | 1,000 t/s | 250K TPM, 1K RPM |
| GPT OSS Safeguard 20B | $0.075 input / $0.30 output | 1,000 t/s | 150K TPM, 1K RPM |
| Qwen3.8-27B (preview) | $0.80 input / $4.00 output | 450 t/s | 250K TPM, 1K RPM |
| whisper-large-v3 | $0.111 per audio hour | n/a | 200K ASH, 300 RPM |
| whisper-large-v3-turbo | $0.04 per audio hour | n/a | 400K ASH, 400 RPM |
| Orpheus V1 English | $22 per 1M characters | n/a | 50K TPM, 250 RPM |
| Llama 3.1 8B (Enterprise) | Contact Sales | 560 t/s | Contact Sales |
| Llama 3.3 70B (Enterprise) | Contact Sales | 280 t/s | Contact Sales |
| MiniMax M2.7 (Enterprise) | Contact Sales | 260 t/s | Contact Sales |
FAQ
Is Groq free to use?
Yes. The free plan covers every listed model with per-model caps of 30 requests per minute, 1,000 requests per day, 8K tokens per minute and 200K tokens per day on the GPT OSS models and Qwen3.8-27B. The developer plan raises those limits and adds Batch and Flex processing.
How much does Groq cost per 1M tokens?
GPT OSS 20B is $0.075 input and $0.30 output, GPT OSS 120B is $0.15 and $0.60, and Qwen3.8-27B is $0.80 and $4.00, per Groq's models page on September 21, 2026. Whisper bills per audio hour instead, at $0.111 for large-v3 and $0.04 for turbo.
Is Groq cheaper than Cerebras?
Per token, yes on GPT OSS 120B: Groq lists $0.15/$0.60 against Cerebras' $0.35/$0.75 per 1M tokens. Per second of wall-clock time it reverses, because Cerebras quotes about 3,000 tokens per second for that model against Groq's 500.
Related reading
Qwen3.8-27B on Cerebras and Groq, cheapest LLM API pricing, OpenRouter alternatives, cheapest GPU cloud for LLM inference