Free LLM APIs for Developers — API Specs, Endpoints & Rate Limits

Developer API specifications for verified free-tier LLM providers. OpenAI-compatible endpoints, authentication methods, rate limits (RPM/RPD/TPD), streaming support, and verification timestamps — all provenance-labeled.
Last verified: Invalid Date

Free LLM APIs: No Credit Card & No Subscription Required

The following providers grant API access with zero payment verification. No credit card, no subscription, and no billing required — just an API key to start building.

Google AI Studio
15 RPM / 1M tokens/day
Gemini 2.5 Flash — no card required
Groq
30 RPM / 1,000,000 TPM
Llama 3.3 70B — no card required
Cerebras
1M tokens/day
Llama 3.3 70B — verified dev tier
OpenRouter
20 RPM / free models
:free models — no card required
Hugging Face
1,000 req/day
Inference API — no card required
Cloudflare Workers AI
5 RPM / 10k/day
Llama 3.1 8B — no card required

API Compatibility Matrix

Quick comparison of endpoint formats, auth methods, and streaming support across all free-tier providers.

ProviderAPI FormatAuthStreamingRate LimitsVerifiedModels
Google AI SDKAPI Key (no CC)✅ SSE15 RPM, 1M tokens/dayPermanent Free3 verified
OpenAI CompatibleAPI Key (no CC)✅ SSE30 RPM, 14,400 req/dayPermanent Free6 verified
SA
SambaNova SystemsAggregator
OpenAI CompatibleAPI Key (no CC)✅ SSE~20 RPM, 200 RPD daily quotaFree Tier3 verified
OpenAI CompatibleAPI Key (no CC)✅ SSE30 RPM, 1M tokens/dayPermanent Free3 verified
OpenAI CompatibleAPI Key (no CC)✅ SSE1,000 promotional API credits for developersTrial4 verified
Custom HTTPAPI Key (no CC)—1 RPS (60 RPM), phone verification requiredUnknown3 verified
Custom HTTPAPI Key (no CC)—10,000 neurons/day free allocationUnknown5 verified
OR
OpenRouterAggregator
OpenAI CompatibleAPI Key (no CC)✅ SSE3 RPM, no credit card required for :free modelsFree Aggregator35 verified
🤗
Hugging FaceAggregator
OpenAI CompatibleAPI Key (no CC)✅ SSE1,000 requests/day via Inference API (serverless)Free Tier3 verified
OpenAI CompatibleAPI Key (no CC)✅ SSEGitHub account required; Copilot limits applyTrial5 verified
OpenAI CompatibleAPI Key (no CC)✅ SSEDeveloper trial quota; varies by modelFree Aggregator5 verified
CH
Chutes.aiAggregator
Custom HTTPAPI Key (no CC)—Free tier with rate limitsUnknown3 verified
MO
ModelScopeAggregator
Custom HTTPAPI Key (no CC)—Free tier with rate limitsUnknown3 verified
OV
OVHcloud AI EndpointsAggregator
Custom HTTPAPI Key (no CC)—Free tier with rate limitsUnknown3 verified

Provider API Specifications

Detailed API specs per provider with documented rate limits, capabilities, and verification status.

Google AI Studio

Provider page →

Google's official AI studio offering Gemini 2.0 Flash, Gemini 2.5 Flash, and Gemma 2 models with generous free quotas.

Verified Offerings

Gemini 2.0 FlashGoogle AI SDK
ReasoningCodingVisionTool Calling
Rate: 15 RPM, 1M tokens/day via Google AI Studio
Verified: Sep 26, 2026(Verified Live)
Credit card: Not required
Gemini 1.5 FlashGoogle AI SDK
ReasoningCodingVisionTool Calling
Rate: 15 RPM, 1.5M tokens/day via Google AI Studio
Verified: Sep 26, 2026(Verified Live)
Credit card: Not required
Rate Limits15 RPM, 1M tokens/day
Context Window1M tokens
TierUnknown
Speed150 tok/s

Ultra-low latency inference via custom LPU chips. Fastest free tier at 330+ tok/s on Llama 3.3 70B.

Verified Offerings

Llama 3.3 70BOpenAI Compatible
ReasoningCodingVisionTool Calling
Rate: 30 RPM, 14,400 requests/day via Groq Console
Verified: Sep 26, 2026(Verified Live)
Credit card: Not required
Llama 3.1 8BOpenAI Compatible
ReasoningCodingVisionTool Calling
Rate: 30 RPM, 14,400 requests/day via Groq Console
Verified: Sep 26, 2026(Verified Live)
Credit card: Not required
Qwen 2.5 Coder 32BOpenAI Compatible
ReasoningCodingVisionTool Calling
Rate: 30 RPM, 14,400 requests/day via Groq Console
Verified: Sep 26, 2026(Verified Live)
Credit card: Not required
Qwen 2.5 72BOpenAI Compatible
ReasoningCodingVisionTool Calling
Rate: 30 RPM, 14,400 requests/day via Groq Console
Verified: Sep 26, 2026(Verified Live)
Credit card: Not required
Rate Limits30 RPM, 14,400 req/day
Context Window128K tokens
TierUnknown
Speed330 tok/s
SA

SambaNova Systems

Aggregator
Provider page →

Reconfigurable dataflow unit (RDU) inference cloud. Sub-second latency on SN40L RDUs with daily free quota.

Verified Offerings

Llama 3.3 70BOpenAI Compatible
ReasoningCodingVisionTool Calling
Rate: ~20 RPM / 200 RPD daily free quota via SN40L RDUs, sub-second latency
Verified: Sep 26, 2026(Documented Free)
Credit card: Not required
Rate Limits~20 RPM, 200 RPD daily quota
Context Window128K tokens
TierUnknown
Speed250 tok/s

Wafer-scale AI inference on custom CS-3 systems. Fastest free tier at 2,100+ tok/s on Llama 3.3 70B.

Verified Offerings

Llama 3.3 70BOpenAI Compatible
ReasoningCodingVisionTool Calling
Rate: 30 RPM, 1M tokens/day via Cerebras Inference API
Verified: Sep 26, 2026(Documented Free)
Credit card: Not required
Llama 3.1 8BOpenAI Compatible
ReasoningCodingVisionTool Calling
Rate: 30 RPM, 1M tokens/day via Cerebras Inference API
Verified: Sep 26, 2026(Documented Free)
Credit card: Not required
Rate Limits30 RPM, 1M tokens/day
Context Window8K tokens
TierUnknown
Speed2100 tok/s

NVIDIA inference microservices with 1,000 free promotional API credits. Enterprise-grade GPU-accelerated inference.

Verified Offerings

Llama 3.3 70BOpenAI Compatible
ReasoningCodingVisionTool Calling
Rate: 1,000 free promotional API credits for developers; standard rate limits apply after trial
Verified: Sep 26, 2026(Documented Free)
Credit card: Not required
Rate Limits1,000 promotional API credits for developers
Context Window128K tokens
TierUnknown
Speed180 tok/s

European open-source LLM platform with MoE architecture. Free tier requires phone verification.

No verified free offerings for this provider.

Rate Limits1 RPS (60 RPM), phone verification required
Context Window128K tokens
TierUnknown
Speed200 tok/s
CF

Cloudflare Workers AI

Provider page →

Edge AI inference on Cloudflare's global network. 10,000 neurons/day free for text, audio, vision, and embeddings.

No verified free offerings for this provider.

Rate Limits10,000 neurons/day free allocation
Context Window128K tokens
TierUnknown
Speed80 tok/s
OR

OpenRouter

Aggregator
Provider page →

Multi-provider routing layer aggregating 35+ free model endpoints. :free models available with rate limits.

Verified Offerings

Meta Llama 3.3 70BOpenAI Compatible
ReasoningCodingVisionTool Calling
Rate: 3 RPM, no credit card required for :free models
Verified: Sep 26, 2026(Documented Free)
Credit card: Not required
Qwen 2.5 72BOpenAI Compatible
ReasoningCodingVisionTool Calling
Rate: 3 RPM, no credit card required for :free models
Verified: Sep 26, 2026(Documented Free)
Credit card: Not required
Meta Llama 3.1 8BOpenAI Compatible
ReasoningCodingVisionTool Calling
Rate: 3 RPM, no credit card required for :free models
Verified: Sep 26, 2026(Documented Free)
Credit card: Not required
Rate Limits3 RPM, no credit card required for :free models
Context Window164K tokens
TierUnknown
Speed90 tok/s
🤗

Hugging Face

Aggregator
Provider page →

Open-source model hub with serverless Inference API. 1,000 requests/day free on serverless endpoints.

Verified Offerings

Llama 3.1 8B InstructOpenAI Compatible
ReasoningCodingVisionTool Calling
Rate: 1,000 requests/day via Inference API (serverless)
Verified: Sep 26, 2026(Verified Live)
Credit card: Not required
Llama 3.3 70B InstructOpenAI Compatible
ReasoningCodingVisionTool Calling
Rate: 500 requests/day via Inference API (serverless, rate-limited)
Verified: Sep 26, 2026(Documented Free)
Credit card: Not required
Rate Limits1,000 requests/day via Inference API (serverless)
Context Window128K tokens
TierUnknown
Speed30 tok/s

GitHub Models

Provider page →

Free AI models via GitHub Models with GitHub account. GPT-4o mini, Llama 3.3, Phi-4 available with Copilot rate limits.

Verified Offerings

GPT-4o miniOpenAI Compatible
ReasoningCodingVisionTool Calling
Rate: GitHub account required; Copilot limits apply (15 RPM, 150 RPD)
Verified: Sep 26, 2026(Documented Free)
Credit card: Not required
Llama 3.3 70BOpenAI Compatible
ReasoningCodingVisionTool Calling
Rate: GitHub account required; Copilot limits apply (15 RPM, 150 RPD)
Verified: Sep 26, 2026(Documented Free)
Credit card: Not required
Phi-4 (14B)OpenAI Compatible
ReasoningCodingVisionTool Calling
Rate: GitHub account required; Copilot limits apply (15 RPM, 150 RPD)
Verified: Sep 26, 2026(Documented Free)
Credit card: Not required
Rate LimitsGitHub account required; Copilot limits apply
Context Window128K tokens
TierUnknown
Speed60 tok/s
KI

Kilo Code / Kilo Gateway

Provider page →

Free AI model gateway routing to frontier models including Qwen 2.5 Coder, DeepSeek Coder, and MiMo V2.5 with zero-cost access.

Verified Offerings

Qwen 2.5 Coder 32BOpenAI Compatible
ReasoningCodingVisionTool Calling
Rate: Developer trial quota; varies by model
Verified: Sep 26, 2026(Documented Free)
Credit card: Not required
DeepSeek Coder V2 (Distill Qwen 32B)OpenAI Compatible
ReasoningCodingVisionTool Calling
Rate: Developer trial quota; varies by model
Verified: Sep 26, 2026(Documented Free)
Credit card: Not required
Rate LimitsDeveloper trial quota; varies by model
Context Window128K tokens
TierUnknown
Speed150 tok/s
CH

Chutes.ai

Aggregator

Community-hosted inference with free tiers. Fast deployment for open-source LLMs with competitive rates.

No verified free offerings for this provider.

Rate LimitsFree tier with rate limits
Context Window128K tokens
TierUnknown
Speed100 tok/s
MO

ModelScope

Aggregator

Chinese AI model hub with free inference endpoints for text, code, and vision models.

No verified free offerings for this provider.

Rate LimitsFree tier with rate limits
Context Window128K tokens
TierUnknown
Speed80 tok/s
OV

OVHcloud AI Endpoints

Aggregator

European sovereign AI endpoints with free tier. GDPR-compliant inference on OVHcloud infrastructure.

No verified free offerings for this provider.

Rate LimitsFree tier with rate limits
Context Window128K tokens
TierUnknown
Speed50 tok/s

Quick Start: OpenAI-Compatible Free APIs

The fastest route to free LLM inference is using OpenAI-compatible endpoints. Below are tested curl examples for the most reliable free providers:

Groq (OpenAI-compatible)

curl https://api.groq.com/openai/v1/chat/completions \
  -H "Authorization: Bearer $GROQ_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "llama-3.3-70b-versatile",
    "messages": [{"role": "user", "content": "Explain quantum computing in 2 sentences."}],
    "stream": true
  }'

OpenRouter (OpenAI-compatible)

curl https://openrouter.ai/api/v1/chat/completions \
  -H "Authorization: Bearer $OPENROUTER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "meta-llama/llama-3.3-70b:free",
    "messages": [{"role": "user", "content": "Explain quantum computing in 2 sentences."}],
    "stream": true
  }'

Together.ai (OpenAI-compatible)

curl https://api.together.ai/v1/chat/completions \
  -H "Authorization: Bearer $TOGETHER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "meta-llama/Llama-3.3-70B-Instruct-Turbo",
    "messages": [{"role": "user", "content": "Explain quantum computing in 2 sentences."}],
    "stream": true
  }'

All providers above support no-credit-card free tiers with varying rate limits. See individual provider sections above for exact quotas, last verified dates, and signup links.

Need hosting recommendations?

Free APIs are great for prototyping. When you need higher rate limits or dedicated infrastructure, compare cloud GPU pricing.

Compare Cloud GPU Pricing