ReasoningDenseContext: 125K
GPT-4o
Comprehensive deployment profile and benchmark telemetry for GPT-4o. Self-hosting memory footprints, verified token economics, and direct API endpoints.
Verified Engineering Benchmarks
SWE-bench
49%
LiveCodeBench
42%
MATH-500
75.5%
MMLU
86.4%
VRAM Requirements & Sizing
🔗 Quick Actions
🖥️ Compatible GPUs & Self-Host Pricing
Deterministic VRAM math from entity graph. Green = fits in single GPU.
| GPU | VRAM | FP16 | FP8 | INT4 | Calculator Link |
|---|---|---|---|---|---|
| H100 SXM5 | 80 GB | OOM | OOM | OOM | Pre-filled → |
| H200 | 141 GB | OOM | OOM | OOM | Pre-filled → |
| B200 | 192 GB | OOM | OOM | OOM | Pre-filled → |
| A100 80GB | 80 GB | OOM | OOM | OOM | Pre-filled → |
| L40S | 48 GB | OOM | OOM | OOM | Pre-filled → |
| RTX 4090 | 24 GB | OOM | OOM | OOM | Pre-filled → |
Active API Providers & Free Tiers
| Provider | Input / 1M | Output / 1M | Speed | Free Tier Status |
|---|---|---|---|---|
| OpenAI | $2.50 | $10.00 | 80 tok/s | Paid API Only |
| OpenRouter | $2.50 | $10.00 | 70 tok/s | Paid API Only |
| NVIDIA NIM | $0.10 | $0.15 | 500 tok/s | Paid API Only |
| GitHub Models | $0.15 | $0.20 | 30 tok/s | ⚡ 15 RPM / 150 RPD |