DeepInfra — Serverless Open-Source Inference
Workload Suitability Matrix: Multi-Node Training: Serverless architecture not designed for distributed training; auto-scales from zero to production without capacity planning. Spot Inference Prototypi...
Hosted Models & Token Rates
| MODEL | TOKENS/SEC | INPUT /1M | OUTPUT /1M | FREE TIER | CONTEXT |
|---|---|---|---|---|---|
DeepSeek R1 Distill 70B | 100 tok/s | $0.01 | $0.01 | — | 128K |
DeepSeek R1 Distill Qwen 32B | 100 tok/s | $0.01 | $0.01 | 20 RPM | 128K |
Llama 3.3 70B Instruct | 500 tok/s | $0.00 | $0.00 | 30 RPM / 14,400 RPD | 128K |
Llama 3.1 8B Instruct | 800 tok/s | $0.01 | $0.01 | 30 RPM / 14,400 RPD | 128K |
Llama 3.2 3B Instruct | 1,200 tok/s | $0.02 | $0.02 | 30 RPM / 14,400 RPD | 128K |
Llama 3.2 1B Instruct | 1,500 tok/s | $0.01 | $0.01 | 30 RPM / 14,400 RPD | 128K |
Llama 3.1 70B Instruct | 90 tok/s | $0.35 | $0.35 | — | 128K |
Llama 3.1 405B Instruct | 14 tok/s | $2.50 | $2.50 | — | 128K |
Qwen 2.5 Coder 32B | 500 tok/s | $0.10 | $0.15 | 20 RPM | 128K |
Qwen 2.5 Coder 14B | 120 tok/s | $0.10 | $0.10 | 20 RPM | 128K |
Qwen 2.5 Coder 7B | 300 tok/s | $0.05 | $0.05 | 20 RPM | 128K |
Qwen 2.5 72B Instruct | 50 tok/s | $0.35 | $0.40 | — | 128K |
Qwen 2.5 14B Instruct | 120 tok/s | $0.10 | $0.10 | 20 RPM | 128K |
Qwen 2.5 7B Instruct | 300 tok/s | $0.04 | $0.04 | 20 RPM | 128K |
Qwen 2.5 VL 72B | 45 tok/s | $0.35 | $0.40 | — | 128K |
Mistral Small v2409 24B | 70 tok/s | $0.10 | $0.10 | — | 128K |
Mistral NeMo 12B | 120 tok/s | $0.07 | $0.09 | — | 128K |
Mixtral 8x22B Instruct | 25 tok/s | $0.50 | $0.50 | — | 65K |
Gemma 2 27B | 250 tok/s | $0.10 | $0.10 | 30 RPM / 14,400 RPD | 8K |
Gemma 2 9B | 550 tok/s | $0.05 | $0.05 | 30 RPM / 14,400 RPD | 8K |
Gemma 3 12B | 400 tok/s | $0.07 | $0.07 | 30 RPM / 14,400 RPD | 128K |
Phi-4 14B | 100 tok/s | $0.10 | $0.14 | — | 16K |
SmolLM2 1.7B | 400 tok/s | $0.05 | $0.05 | — | 128K |
Mistral 7B v0.3 | 250 tok/s | $0.06 | $0.06 | — | 32K |
Competitor Comparison
| PROVIDER | MIN PRICE / HR | AVG TOKEN RATE | FREE TIER | BILLING | KEY ADVANTAGE |
|---|---|---|---|---|---|
| DeepInfra | — | $0.00/1M | Generous free tier with rate l... | per-token | ★ You are here |
| OpenRouter | — | $0.00/1M | :free models available with ra... | per-token | Multi-Provider Routing Layer |
Technical Nuances & Editorial Analysis
Workload Suitability Matrix: Multi-Node Training: Serverless architecture not designed for distributed training; auto-scales from zero to production without capacity planning. Spot Inference Prototyping: Lowest input and output token pricing in the market; supports 25+ model variants across DeepSeek, Llama, Qwen, Mistral, Gemma families; serverless auto-scales without capacity planning. Persistent Production API: Fully OpenAI-compatible API; supports streaming, batch processing, and fine-tuning APIs; ideal for developers switching between models without infrastructure overhead.
Billing Granularity
per-token
How charges are calculated
Hidden Costs
Pay-per-token model with no minimum commitments; serverless auto-scales
Watch out for these fees
Free Tier
Generous free tier with rate limits on select models
Credit requirements and caps
Best Use Case
Best for API-based LLM inference
Ideal workload profile