⚡Under $0.50/hr🧠VRAM Estimator⚖Compare GPUs🎁Free LLM APIs🎯Model Index
LLM APIper-token Live

Groq — LPU Inference Engine

Workload Suitability Matrix: Multi-Node Training: LPU architecture not optimized for distributed training; best suited for inference workloads. Spot Inference Prototyping: 30 RPM / 14,400 RPD free tie...

⚡ Free Tier Summary — Groq
Llama 3.3 70BNo Card
30 RPM, 14,400 requests/day via Groq Console
Llama 3.1 8BNo Card
30 RPM, 14,400 requests/day via Groq Console
Qwen 2.5 Coder 32BNo Card
30 RPM, 14,400 requests/day via Groq Console
Qwen 2.5 72BNo Card
30 RPM, 14,400 requests/day via Groq Console
Meta Llama 3.3 70BNo Card
3 RPM, no credit card required for :free models
Meta Llama 3.1 8BNo Card
3 RPM, no credit card required for :free models
Llama 3.1 8B InstructNo Card
1,000 requests/day via Inference API (serverless)
Llama 3.3 70B InstructNo Card
500 requests/day via Inference API (serverless, rate-limited)
Llama 3.3 70BNo Card
10,000 neurons/day free allocation
Llama 3.1 8BNo Card
10,000 neurons/day free allocation

Hosted Models & Token Rates

MODELTOKENS/SECINPUT /1MOUTPUT /1MFREE TIERCONTEXT
Llama 3.3 70B Instruct
500 tok/s$0.00$0.0030 RPM / 14,400 RPD128K
Llama 3.1 8B Instruct
800 tok/s$0.01$0.0130 RPM / 14,400 RPD128K
Llama 3.2 3B Instruct
1,200 tok/s$0.02$0.0230 RPM / 14,400 RPD128K
Llama 3.2 1B Instruct
1,500 tok/s$0.01$0.0130 RPM / 14,400 RPD128K
Gemma 2 27B
250 tok/s$0.10$0.1030 RPM / 14,400 RPD8K
Gemma 2 9B
550 tok/s$0.05$0.0530 RPM / 14,400 RPD8K
Gemma 3 12B
400 tok/s$0.07$0.0730 RPM / 14,400 RPD128K

Competitor Comparison

PROVIDERMIN PRICE / HRAVG TOKEN RATEFREE TIERBILLINGKEY ADVANTAGE
Groq—$0.00/1M30 requests per minute (RPM) f...per-token★ You are here
Cerebras—$0.00/1MLimited free tier available; c...per-tokenWafer-Scale AI Inference
DeepInfra—$0.00/1MGenerous free tier with rate l...per-tokenServerless Open-Source Inference

Technical Nuances & Editorial Analysis

Workload Suitability Matrix: Multi-Node Training: LPU architecture not optimized for distributed training; best suited for inference workloads. Spot Inference Prototyping: 30 RPM / 14,400 RPD free tier requires no credit card; API fully OpenAI-compatible for drop-in replacement; token generation 280-800 tok/s. Persistent Production API: Model catalog smaller than competitors; free tier available with rate limits; paid tier for higher throughput; ideal for real-time streaming, voice assistants, and latency-sensitive deployments.

Billing Granularity

per-token

How charges are calculated

Hidden Costs

No credit card required for free tier; paid tier requires API key

Watch out for these fees

Free Tier

30 requests per minute (RPM) free tier; no credit card required

Credit requirements and caps

Best Use Case

Best for API-based LLM inference

Ideal workload profile

Frequently Asked Questions

What is the cheapest instance/model on Groq?▾
The lowest input token rate on Groq is $0.00/1M tokens via SambaNova, Cloudflare Workers AI, Groq.
Does Groq offer a free tier without a credit card?▾
30 requests per minute (RPM) free tier; no credit card required
How does Groq compare to its cheaper alternatives?▾
Groq differentiates through lpu inference engine. Token rates range from $0.00/1M input, competitive with the broader API market.
Data Freshness: Verified via Official Provider APIs & Documentation | Refreshed WeeklyPricing sourced from Groq official documentation and public APIs.Methodology →

What should I do next?