GPU Comparison Engine
Select any two GPUs for a side-by-side evaluation of VRAM capacity, memory bandwidth, TFLOPS throughput, and live spot pricing across workload profiles.
Methodology: How GPU Comparison Works
OpenGPU Radar compares GPU instances across four dimensions: Price, Performance, Memory, and Availability. Pricing data is aggregated from Spheron, RunPod, Vast.ai, and Lambda Labs APIs, refreshed every 6 hours.
Cost per 1M Tokens = (Hourly GPU Cost) / (Tokens/Second × 1,000,000) Lower values indicate better economics. Throughput Index = Effective Tokens/Second (vLLM FP8, Llama 70B) Measured end-to-end output throughput including KV-cache retrieval. Example: H100 SXM5 at $1.89/hr delivers ~120 tok/s → $1.89 / (120 × 1M) = $0.000158 per 1M tokens
Hardware Requirements by Workload
Minimum VRAM and GPU specs required for each workload category. Determined via entity-graph VRAM calculations.
| Workload | Min GPU | Min VRAM | Precision | Recommended GPU |
|---|---|---|---|---|
| 70B LLM Inference | H100 SXM5 | 80 GB | FP8 | H200 or B200 |
| 30B LLM Inference | L40S | 48 GB | FP8 | A100 80GB |
| 7B LLM Inference | RTX 4090 | 24 GB | FP16 | RTX 4090 |
| Fine-Tuning (LoRA) | A100 80GB | 80 GB | FP16 | H100 × 2 |
| Pre-Training | B200 | 192 GB | FP8 | B200 × 8 |
| Vision Model (72B) | H200 | 141 GB | FP8 | B200 |
Workload Suitability Profiles
Interconnect and RDMA availability for distributed workloads
Memory bandwidth and tensor-core throughput trade-offs
Storage volume costs and SLA guarantees for persistent APIs
Interruptible risk profile vs dedicated enterprise reliability
Spot pricing savings vs guaranteed uptime SLAs
Preemption risk and egress bandwidth charges per workload