Google AI Studio — Gemini API Platform
Workload Suitability Matrix: Multi-Node Training: Native multimodal attention processes image patches, video frames, and text tokens through shared attention mechanism; 1M token context window via spa...
Hosted Models & Token Rates
| MODEL | TOKENS/SEC | INPUT /1M | OUTPUT /1M | FREE TIER | CONTEXT |
|---|---|---|---|---|---|
Gemini 2.5 Flash | 150 tok/s | $0.10 | $0.40 | 15 RPM / 1M tokens/day | 1M |
Gemini 2.0 Flash | 130 tok/s | $0.07 | $0.30 | 15 RPM / 1M tokens/day | 1M |
Gemini 1.5 Flash | 15 tok/s | $0.00 | $0.00 | 1.049M |
Technical Nuances & Editorial Analysis
Workload Suitability Matrix: Multi-Node Training: Native multimodal attention processes image patches, video frames, and text tokens through shared attention mechanism; 1M token context window via sparse attention. Spot Inference Prototyping: 15 RPM / 1,500 RPD free tier on Google AI Studio; no credit card required; prompt caching reduces token costs by up to 90% on repeated context prefixes. Persistent Production API: Paid tiers start at $0.075/$0.30 per million input/output tokens for Gemini Flash; Gemini 1.5 Pro costs $1.25/$5.00 per million tokens; integrates with Google Cloud ecosystem.
Billing Granularity
per-token
How charges are calculated
Hidden Costs
Generous free tier; prompt caching discounts reduce costs by up to 90%
Watch out for these fees
Free Tier
15 RPM free tier for Gemini 2.0 Flash and Pro; no credit card required
Credit requirements and caps
Best Use Case
Best for API-based LLM inference
Ideal workload profile