H100 vs A100 vs L40S for LLM Inference: Practical GPU Selection Guide
When to use H100, A100, or L40S for LLM inference. VRAM fit, site throughput records, and honest cost-per-token math for 7B through 70B models.
Direct answer
The GPU choice for LLM inference depends first on VRAM fit โ and the fit changes with precision and context:
| Model | FP16 weights | H100 (80GB) | A100 (80GB) | L40S (48GB) |
|---|---|---|---|---|
| Llama 3.1 8B | 16 GB | single GPU (FP8: 8 GB) | single GPU | single GPU (FP8: 8 GB) |
| Mistral 7B | 14 GB | single GPU | single GPU | single GPU |
| Qwen 2.5 32B | 65 GB (FP16) / 41 GB (FP8@32K) | FP8 single GPU | FP8 single GPU | FP8 (41 GB) only |
| Llama 3.3 70B | 141 GB (FP16) / 81 GB (FP8@32K) | FP8 needs TP=2 or short context | FP8 no (no native FP8) | FP16/FP8 no ยท INT4 (43 GB) fits |
Choose H100 for FP8 serving of 70B-class models when cost-per-token matters. Choose A100 if you need FP16/BF16 and already operate it โ it has no native FP8 path. Choose L40S for 7Bโ32B models at FP8, or 70B at INT4, where the cheapest per-token rate wins.
The GPU comparison tool has live pricing for this decision.
VRAM is the first filter
Parameter count gives the weights floor; service VRAM adds KV cache, CUDA context, and runtime overhead (canonical engine: (weights + KV + 1.6 GB) ร 1.10):
| Model | Params (B) | FP16 (weights) | FP8 (weights) | INT4 (weights) |
|---|---|---|---|---|
| Llama 3.1 8B | 8.03 | 16 GB | 8 GB | 4 GB |
| Mistral 7B | 7.0 | 14 GB | 7 GB | 3.5 GB |
| Qwen 2.5 32B | 32.5 | 65 GB | 33 GB | 16 GB |
| Llama 3.3 70B | 70.6 | 141 GB | 71 GB | 35 GB |
| Qwen 2.5 72B | 72.7 | 145 GB | 73 GB | 36 GB |
These rows are weights only. Service totals at 32K context: 8B FP8 = 12.8 GB, 70B FP8 = 81.3 GB, Qwen 32B FP8 = 41.1 GB. At 128K, KV cache adds a lot: for 70B it's ~7 GB with the registry's 28-layer preset or ~20 GB with the 80-layer preset โ the repo's layer counts for 70B models disagree (28 vs 80), so size KV as a 7โ20 GB range until that data is reconciled (87โ101 GB total FP8 @128K).
H100 SXM5: The FP8 sweet spot
| Spec | Value |
|---|---|
| VRAM | 80 GB HBM3 |
| FP8 TFLOPS | 1,979 (2ร the FP16 convention) |
| 70B FP8 throughput | ~120 tok/s (site record, batch=1) |
| Observed price | $1.89/hr (Vast.ai, 2026-10-03) |
The H100 is the right choice when:
- You're serving 70B-class models in FP8 โ note the fit boundary: 81 GB at 32K > 80 GB, so single-GPU H100 only works at short context (โ4K: 79.5 GB); 32K needs TP=2
- You want native FP8 Tensor Cores (1,979 FP8 TFLOPS)
- You need NVLink 4.0 for multi-GPU tensor parallelism
- Cost-per-token: $1.89/hr รท 120 tok/s = $4.38/M tokens at batch=1 โ a latency-oriented figure; production batching cuts this substantially
Check live H100 rates โ observed spread is $1.89โ$3.49/hr across providers (~85%).
A100 SXM4: FP16/BF16 legacy
| Spec | Value |
|---|---|
| VRAM | 80 GB HBM2e (2.0 TB/s) |
| FP16 TFLOPS | 312 |
| FP8 path | none in silicon โ repo lists 624 as a derived/emulated figure, not a vendor spec |
| 70B throughput | ~45 tok/s (site record: 70B INT8, batch=1, 2-GPU TP) |
| Observed price | $1.59/hr (Lambda Labs) |
The A100 is the right choice when:
- You need BF16/FP16 (or INT8) โ not FP8: A100 has no FP8 tensor path, so "FP8 on A100" is emulation and should not be planned as such
- You have existing A100 infrastructure and want to avoid migration
- Your model is โค80 GB at FP16 (single GPU without TP)
- 70B INT8 at TP=2: 2 ร $1.59 = $3.18/hr รท 45 tok/s = $19.7/M at batch=1 โ the site's only A100 70B record
See A100-80GB specs for details.
L40S: Budget inference for smaller models
| Spec | Value |
|---|---|
| VRAM | 48 GB GDDR6 |
| FP8 TFLOPS | 733 |
| Throughput record | ~65 tok/s (8B FP8, batch=1) โ no 70B FP8 record exists |
| Observed price | $0.69/hr (Vast.ai) |
The L40S is the right choice when:
- Your model fits in 48 GB: 7Bโ32B at FP8, 8B at FP16, 70B at INT4 (42.5 GB โ fits, 5.5 GB headroom)
- Cost-per-token is the priority: 8B FP8 at $0.69/hr รท 65 tok/s = $2.95/M (batch=1)
- You need a standard PCIe datacenter card (no NVLink โ multi-GPU is PCIe-bound)
- 70B serving must stay single-GPU โ INT4 is the only precision that fits; there is no published throughput record for that config, so measure it yourself
The old claim here โ "L40S delivers roughly the same 70B FP8 throughput as H100" โ was wrong and has been removed. A 70B FP8 model (81 GB) cannot run on an L40S (48 GB) at all; the L40S's only 70B-class record in the spec database is A100's 45 tok/s INT8 record, which belongs to a different GPU. See L40S specs.
When VRAM is the binding constraint
At 128K context, Llama 3.3 70B needs 87โ101 GB at FP8 (KV layer-count range above). None of H100, A100, or L40S fits it single-GPU. Your options:
- H200 SXM5 (141 GB) โ single-GPU FP8 at 128K
- 2ร H100 with tensor parallelism โ NVLink overhead, ~2ร hourly cost ($3.78/hr at Vast.ai rates)
- INT4 quantization โ 42.5 GB service total: fits A100 (80), H100 (80), and L40S (48) single-GPU
The H100 vs H200 ROI analysis walks through the cost decision for 128K deployments.
Quantization changes the equation
INT4 compresses weights 4ร relative to FP16 (weights-only figures; service totals add KV + ~7.5% overhead):
| Model | FP16 (weights) | INT4 (weights) | Service INT4 total | Fits on L40S (48GB)? |
|---|---|---|---|---|
| Llama 3.1 8B | 16 GB | 4 GB | ~5 GB | yes |
| Mistral 7B | 14 GB | 3.5 GB | ~4.5 GB | yes |
| Qwen 2.5 32B | 65 GB | 16 GB | ~18 GB | yes |
| Llama 3.3 70B | 141 GB | 35 GB | 42.5 GB | yes (5.5 GB headroom) |
| Qwen 2.5 72B | 145 GB | 36 GB | ~40 GB | yes |
INT4 quality impact is workload-specific โ the quantization guide covers how to evaluate it.
Decision tree
- Model โค 30 GB service VRAM at your precision? โ L40S (cheapest per-token: $2.95/M at 8B FP8 batch=1)
- 70B-class at FP8? โ H100 (short context, single GPU) or 2ร H100 TP (32K); H200 for 128K
- FP16/BF16 required and A100 already on hand? โ A100 โ but do not plan "FP8 on A100"
- 70B on a budget, batch=1 acceptable? โ L40S at INT4 (42.5 GB) โ benchmark your own throughput, no record exists
- VRAM exceeds 80 GB at FP8 (128K)? โ H200 or multi-GPU H100
Limitations
- Throughput figures are site GPU spec records (batch=1, named configs); combinations without records are marked as such rather than projected
- VRAM figures exclude PagedAttention savings
- GPU prices are [OBSERVED] as of 2026-10-05 and fluctuate
- A100 has no native FP8 โ any FP8 number for it is derived, not measured
- 70B KV cache is a 7โ20 GB range pending the repo's layer-count reconciliation
Related resources
- H100 SXM5 GPU Specs โ live pricing
- A100 80GB GPU Specs โ FP16 inference
- L40S GPU Specs โ budget FP8 inference
- GPU comparison โ search by model size
- VRAM calculator โ model-specific memory math