⚡Under $0.50/hr🧠VRAM Estimator⚖Compare GPUs🎁Free LLM APIs🎯Model Index

LLM Model Registry 48 Models • 21 Providers

Real-time token pricing, verified free rate limits, and minimum VRAM footprints.

🟢 Live Free Tier Health: 24 verified free model endpoints active today. No credit cards required.
MODEL & FAMILYCONTEXTBEST API RATE ($/1M)SPEEDFREE TIERMIN VRAMACTION

DeepSeek R1 ↗

DeepSeekMoEReasoning⚡ Tools/JSON
128K$0.00/0.00500 tok/s—670GB ↗

Llama 3.3 70B Instruct ↗

MetaDenseReasoning⚡ Tools/JSON
128K$0.00/0.00500 tok/sGroq (30 RPM / 14,400 RPD)No Card40GB ↗

FLUX.1 [dev] ↗

Black Forest LabsDenseMultimodal
—$0.00/0.005 tok/s—12GB ↗

FLUX.1 [schnell] ↗

Black Forest LabsDenseMultimodal
—$0.00/0.006 tok/sTogether AI (Free tier)No Card12GB ↗

MiMo V2.5 (Free) ↗

Kilo GatewayProprietaryCoding⚡ Tools/JSON
128K$0.00/0.00120 tok/sKilo Gateway (Daily Free Tier)No CardClosed

Ling 3.0 Flash VL (Free) ↗

Kilo GatewayProprietaryMultimodal⚡ Tools/JSON
128K$0.00/0.00180 tok/sKilo Gateway (Free in Kilo Code)No CardClosed

GLM-4 Flash ↗

Zhipu AIProprietaryEdge
128K$0.00/0.00100 tok/sZ.ai (Permanently Free)No CardClosed

Gemini 1.5 Flash ↗

GoogleProprietaryMultimodal
1.049M Context$0.00/0.0015 tok/sGoogle AI Studio ()No Card—GB ↗

Codestral 22B ↗

Mistral AIProprietaryCoding
128K$0.00/0.001 tok/sMistral AI (La Plateforme) ()No Card22GB ↗

DeepSeek R1 Distill 70B ↗

DeepSeekDenseReasoning
128K$0.01/0.01100 tok/s—40GB ↗

DeepSeek R1 Distill Qwen 32B ↗

DeepSeekDenseReasoning
128K$0.01/0.01100 tok/sOpenRouter (20 RPM)No Card20GB ↗

Llama 3.1 8B Instruct ↗

MetaDenseCoding⚡ Tools/JSON
128K$0.01/0.01800 tok/sGroq (30 RPM / 14,400 RPD)No Card6GB ↗

Llama 3.2 1B Instruct ↗

MetaDenseEdge⚡ Tools/JSON
128K$0.01/0.011,500 tok/sGroq (30 RPM / 14,400 RPD)No Card2GB ↗

Llama 3.2 3B Instruct ↗

MetaDenseEdge⚡ Tools/JSON
128K$0.02/0.021,200 tok/sGroq (30 RPM / 14,400 RPD)No Card3GB ↗

Qwen 2.5 7B Instruct ↗

AlibabaDenseEdge⚡ Tools/JSON
128K$0.04/0.04300 tok/sOpenRouter (20 RPM)No Card5GB ↗

Qwen 2.5 Coder 7B ↗

AlibabaDenseCoding⚡ Tools/JSON
128K$0.05/0.05300 tok/sOpenRouter (20 RPM)No Card5GB ↗

Gemma 2 9B ↗

Google DeepMindDenseEdge⚡ Tools/JSON
8K$0.05/0.05550 tok/sGroq (30 RPM / 14,400 RPD)No Card6GB ↗

SmolLM2 1.7B ↗

OpenAIDenseEdge⚡ Tools/JSON
128K$0.05/0.05400 tok/s—2GB ↗

Mistral 7B v0.3 ↗

Mistral AIDenseEdge⚡ Tools/JSON
32K$0.06/0.06250 tok/s—5GB ↗

Mistral NeMo 12B ↗

Mistral AIDenseEdge⚡ Tools/JSON
128K$0.07/0.09120 tok/s—8GB ↗

Gemma 3 12B ↗

Google DeepMindDenseGeneral⚡ Tools/JSON
128K$0.07/0.07400 tok/sGroq (30 RPM / 14,400 RPD)No Card8GB ↗

Gemini 2.0 Flash ↗

Google DeepMindProprietaryEdge⚡ Tools/JSON
1M Context$0.07/0.30130 tok/sGoogle AI Studio (15 RPM / 1M tokens/day)No CardClosed

Qwen 2.5 Coder 32B ↗

AlibabaDenseCoding⚡ Tools/JSON
128K$0.10/0.15500 tok/sOpenRouter (20 RPM)No Card20GB ↗

Qwen 2.5 Coder 14B ↗

AlibabaDenseCoding⚡ Tools/JSON
128K$0.10/0.10120 tok/sOpenRouter (20 RPM)No Card10GB ↗

Qwen 2.5 14B Instruct ↗

AlibabaDenseGeneral⚡ Tools/JSON
128K$0.10/0.10120 tok/sOpenRouter (20 RPM)No Card10GB ↗

Mistral Small v2409 24B ↗

Mistral AIDenseCoding
128K$0.10/0.1070 tok/s—14GB ↗

Gemma 2 27B ↗

Google DeepMindDenseGeneral⚡ Tools/JSON
8K$0.10/0.10250 tok/sGroq (30 RPM / 14,400 RPD)No Card16GB ↗

Phi-4 14B ↗

OpenAIDenseEdge
16K$0.10/0.14100 tok/s—10GB ↗

GPT-4o ↗

OpenAIProprietaryReasoning⚡ Tools/JSON
128K$0.10/0.15500 tok/s—Closed

Gemini 2.5 Flash ↗

Google DeepMindProprietaryGeneral⚡ Tools/JSON
1M Context$0.10/0.40150 tok/sGoogle AI Studio (15 RPM / 1M tokens/day)No CardClosed

GPT-4.1 Nano ↗

OpenAIProprietaryEdge⚡ Tools/JSON
128K$0.10/0.40200 tok/s—Closed

Whisper Large v3 Turbo ↗

OpenAIDenseMultimodal
—$0.10/0.10100 tok/s—4GB ↗

Nemotron-4 340B Instruct ↗

NVIDIAMoEMultimodal⚡ Tools/JSON
128K$0.10/0.15500 tok/s—70GB ↗

DeepSeek V3 ↗

DeepSeekMoEReasoning
128K$0.14/0.2835 tok/s—670GB ↗

GPT-4o Mini ↗

OpenAIProprietaryEdge⚡ Tools/JSON
128K$0.15/0.60150 tok/s—Closed

Qwen 2.5 VL 7B ↗

AlibabaDenseMultimodal⚡ Tools/JSON
128K$0.20/0.20150 tok/s—5GB ↗

Mistral Large ↗

Mistral AIProprietaryReasoning
128K$0.20/0.6030 tok/s—37GB ↗

Codestral 22B ↗

Mistral AIDenseCoding
32K$0.30/0.9055 tok/s—14GB ↗

Llama 3.1 70B Instruct ↗

MetaDenseReasoning
128K$0.35/0.3590 tok/s—40GB ↗

Qwen 2.5 72B Instruct ↗

AlibabaDenseReasoning
128K$0.35/0.4050 tok/s—40GB ↗

Qwen 2.5 VL 72B ↗

AlibabaDenseMultimodal
128K$0.35/0.4045 tok/s—40GB ↗

Mixtral 8x22B Instruct ↗

Mistral AIMoEReasoning
65K$0.50/0.5025 tok/s—80GB ↗

Claude 3.5 Haiku ↗

AnthropicProprietaryEdge⚡ Tools/JSON
200K$0.80/4.00120 tok/s—Closed

o3-mini ↗

OpenAIProprietaryReasoning
128K$1.10/4.40100 tok/s—Closed

Llama 3.1 405B Instruct ↗

MetaDenseReasoning
128K$2.50/2.5014 tok/s—230GB ↗

Command R+ ↗

CohereProprietaryReasoning
128K$2.50/10.0025 tok/s—Closed

Claude 3.5 Sonnet ↗

AnthropicProprietaryReasoning
200K$3.00/15.0080 tok/s—Closed

Claude 3 Opus ↗

AnthropicProprietaryReasoning
200K$15.00/75.0030 tok/s—Closed
Data Freshness: Verified via Official Provider APIs & Documentation | Refreshed WeeklyToken pricing sourced from official provider API documentation and OpenRouter catalog.Methodology →