⚡Under $0.50/hr🧠VRAM Estimator⚖Compare GPUs🎁Free LLM APIs🎯Model Index
LLM APIper-token Live

DeepInfra — Serverless Open-Source Inference

Workload Suitability Matrix: Multi-Node Training: Serverless architecture not designed for distributed training; auto-scales from zero to production without capacity planning. Spot Inference Prototypi...

⚡ Free Tier Summary — DeepInfra
Llama 3.3 70BNo Card
30 RPM, 14,400 requests/day via Groq Console
Llama 3.1 8BNo Card
30 RPM, 14,400 requests/day via Groq Console
Qwen 2.5 Coder 32BNo Card
30 RPM, 14,400 requests/day via Groq Console
Qwen 2.5 72BNo Card
30 RPM, 14,400 requests/day via Groq Console
Mistral NeMo 12BNo Card
1 RPS (60 RPM), phone verification required via La Plateforme
Meta Llama 3.3 70BNo Card
3 RPM, no credit card required for :free models
Qwen 2.5 72BNo Card
3 RPM, no credit card required for :free models
Meta Llama 3.1 8BNo Card
3 RPM, no credit card required for :free models
Llama 3.1 8B InstructNo Card
1,000 requests/day via Inference API (serverless)
Llama 3.3 70B InstructNo Card
500 requests/day via Inference API (serverless, rate-limited)
Llama 3.3 70BNo Card
10,000 neurons/day free allocation
Llama 3.1 8BNo Card
10,000 neurons/day free allocation

Hosted Models & Token Rates

MODELTOKENS/SECINPUT /1MOUTPUT /1MFREE TIERCONTEXT
DeepSeek R1 Distill 70B
100 tok/s$0.01$0.01—128K
DeepSeek R1 Distill Qwen 32B
100 tok/s$0.01$0.0120 RPM128K
Llama 3.3 70B Instruct
500 tok/s$0.00$0.0030 RPM / 14,400 RPD128K
Llama 3.1 8B Instruct
800 tok/s$0.01$0.0130 RPM / 14,400 RPD128K
Llama 3.2 3B Instruct
1,200 tok/s$0.02$0.0230 RPM / 14,400 RPD128K
Llama 3.2 1B Instruct
1,500 tok/s$0.01$0.0130 RPM / 14,400 RPD128K
Llama 3.1 70B Instruct
90 tok/s$0.35$0.35—128K
Llama 3.1 405B Instruct
14 tok/s$2.50$2.50—128K
Qwen 2.5 Coder 32B
500 tok/s$0.10$0.1520 RPM128K
Qwen 2.5 Coder 14B
120 tok/s$0.10$0.1020 RPM128K
Qwen 2.5 Coder 7B
300 tok/s$0.05$0.0520 RPM128K
Qwen 2.5 72B Instruct
50 tok/s$0.35$0.40—128K
Qwen 2.5 14B Instruct
120 tok/s$0.10$0.1020 RPM128K
Qwen 2.5 7B Instruct
300 tok/s$0.04$0.0420 RPM128K
Qwen 2.5 VL 72B
45 tok/s$0.35$0.40—128K
Mistral Small v2409 24B
70 tok/s$0.10$0.10—128K
Mistral NeMo 12B
120 tok/s$0.07$0.09—128K
Mixtral 8x22B Instruct
25 tok/s$0.50$0.50—65K
Gemma 2 27B
250 tok/s$0.10$0.1030 RPM / 14,400 RPD8K
Gemma 2 9B
550 tok/s$0.05$0.0530 RPM / 14,400 RPD8K
Gemma 3 12B
400 tok/s$0.07$0.0730 RPM / 14,400 RPD128K
Phi-4 14B
100 tok/s$0.10$0.14—16K
SmolLM2 1.7B
400 tok/s$0.05$0.05—128K
Mistral 7B v0.3
250 tok/s$0.06$0.06—32K

Competitor Comparison

PROVIDERMIN PRICE / HRAVG TOKEN RATEFREE TIERBILLINGKEY ADVANTAGE
DeepInfra—$0.00/1MGenerous free tier with rate l...per-token★ You are here
OpenRouter—$0.00/1M:free models available with ra...per-tokenMulti-Provider Routing Layer

Technical Nuances & Editorial Analysis

Workload Suitability Matrix: Multi-Node Training: Serverless architecture not designed for distributed training; auto-scales from zero to production without capacity planning. Spot Inference Prototyping: Lowest input and output token pricing in the market; supports 25+ model variants across DeepSeek, Llama, Qwen, Mistral, Gemma families; serverless auto-scales without capacity planning. Persistent Production API: Fully OpenAI-compatible API; supports streaming, batch processing, and fine-tuning APIs; ideal for developers switching between models without infrastructure overhead.

Billing Granularity

per-token

How charges are calculated

Hidden Costs

Pay-per-token model with no minimum commitments; serverless auto-scales

Watch out for these fees

Free Tier

Generous free tier with rate limits on select models

Credit requirements and caps

Best Use Case

Best for API-based LLM inference

Ideal workload profile

Frequently Asked Questions

What is the cheapest instance/model on DeepInfra?▾
The lowest input token rate on DeepInfra is $0.00/1M tokens via Cloudflare Workers AI, SambaNova, Groq, DeepInfra, NVIDIA NIM, OpenRouter, Mistral, Together AI.
Does DeepInfra offer a free tier without a credit card?▾
Generous free tier with rate limits on select models
How does DeepInfra compare to its cheaper alternatives?▾
DeepInfra differentiates through serverless open-source inference. Token rates range from $0.00/1M input, competitive with the broader API market.
Data Freshness: Verified via Official Provider APIs & Documentation | Refreshed WeeklyPricing sourced from DeepInfra official documentation and public APIs.Methodology →

Cost Analysis & Deployment Guides

What should I do next?