Cloudflare Workers AI — Edge AI Inference
Workload Suitability Matrix: Multi-Node Training: Workers AI is not optimized for distributed training; designed for inference at the edge. Spot Inference Prototyping: 10,000 neurons/day free allocati...
Hosted Models & Token Rates
| MODEL | TOKENS/SEC | INPUT /1M | OUTPUT /1M | FREE TIER | CONTEXT |
|---|---|---|---|---|---|
DeepSeek R1 Distill 70B | 100 tok/s | $0.01 | $0.01 | — | 128K |
DeepSeek R1 Distill Qwen 32B | 100 tok/s | $0.01 | $0.01 | 20 RPM | 128K |
Llama 3.1 8B Instruct | 800 tok/s | $0.01 | $0.01 | 30 RPM / 14,400 RPD | 128K |
Competitor Comparison
| PROVIDER | MIN PRICE / HR | AVG TOKEN RATE | FREE TIER | BILLING | KEY ADVANTAGE |
|---|---|---|---|---|---|
| Cloudflare Workers AI | — | $0.01/1M | 10,000 neurons/day free alloca... | per-token | ★ You are here |
| Groq | — | $0.00/1M | 30 requests per minute (RPM) f... | per-token | LPU Inference Engine |
| DeepInfra | — | $0.00/1M | Generous free tier with rate l... | per-token | Serverless Open-Source Inference |
| Cerebras | — | $0.00/1M | Limited free tier available; c... | per-token | Wafer-Scale AI Inference |
Technical Nuances & Editorial Analysis
Workload Suitability Matrix: Multi-Node Training: Workers AI is not optimized for distributed training; designed for inference at the edge. Spot Inference Prototyping: 10,000 neurons/day free allocation; global edge network provides low-latency inference; no credit card required. Persistent Production API: Pay-per-use pricing; integrates with Cloudflare ecosystem; supports AI Gateway for rate limiting and caching.
Billing Granularity
per-token
How charges are calculated
Hidden Costs
Free tier has neuron allocation limits; paid tiers per-million tokens
Watch out for these fees
Free Tier
10,000 neurons/day free allocation; no credit card required
Credit requirements and caps
Best Use Case
Best for API-based LLM inference
Ideal workload profile