Benchmarks2026-10-05โ€ขBy Sreeโ€ข5 min read

H100 vs A100 vs L40S for LLM Inference: Practical GPU Selection Guide

When to use H100, A100, or L40S for LLM inference. VRAM fit, site throughput records, and honest cost-per-token math for 7B through 70B models.

Direct answer

The GPU choice for LLM inference depends first on VRAM fit โ€” and the fit changes with precision and context:

ModelFP16 weightsH100 (80GB)A100 (80GB)L40S (48GB)
Llama 3.1 8B16 GBsingle GPU (FP8: 8 GB)single GPUsingle GPU (FP8: 8 GB)
Mistral 7B14 GBsingle GPUsingle GPUsingle GPU
Qwen 2.5 32B65 GB (FP16) / 41 GB (FP8@32K)FP8 single GPUFP8 single GPUFP8 (41 GB) only
Llama 3.3 70B141 GB (FP16) / 81 GB (FP8@32K)FP8 needs TP=2 or short contextFP8 no (no native FP8)FP16/FP8 no ยท INT4 (43 GB) fits

Choose H100 for FP8 serving of 70B-class models when cost-per-token matters. Choose A100 if you need FP16/BF16 and already operate it โ€” it has no native FP8 path. Choose L40S for 7Bโ€“32B models at FP8, or 70B at INT4, where the cheapest per-token rate wins.

The GPU comparison tool has live pricing for this decision.

VRAM is the first filter

Parameter count gives the weights floor; service VRAM adds KV cache, CUDA context, and runtime overhead (canonical engine: (weights + KV + 1.6 GB) ร— 1.10):

ModelParams (B)FP16 (weights)FP8 (weights)INT4 (weights)
Llama 3.1 8B8.0316 GB8 GB4 GB
Mistral 7B7.014 GB7 GB3.5 GB
Qwen 2.5 32B32.565 GB33 GB16 GB
Llama 3.3 70B70.6141 GB71 GB35 GB
Qwen 2.5 72B72.7145 GB73 GB36 GB

These rows are weights only. Service totals at 32K context: 8B FP8 = 12.8 GB, 70B FP8 = 81.3 GB, Qwen 32B FP8 = 41.1 GB. At 128K, KV cache adds a lot: for 70B it's ~7 GB with the registry's 28-layer preset or ~20 GB with the 80-layer preset โ€” the repo's layer counts for 70B models disagree (28 vs 80), so size KV as a 7โ€“20 GB range until that data is reconciled (87โ€“101 GB total FP8 @128K).

H100 SXM5: The FP8 sweet spot

SpecValue
VRAM80 GB HBM3
FP8 TFLOPS1,979 (2ร— the FP16 convention)
70B FP8 throughput~120 tok/s (site record, batch=1)
Observed price$1.89/hr (Vast.ai, 2026-10-03)

The H100 is the right choice when:

  • You're serving 70B-class models in FP8 โ€” note the fit boundary: 81 GB at 32K > 80 GB, so single-GPU H100 only works at short context (โ‰ˆ4K: 79.5 GB); 32K needs TP=2
  • You want native FP8 Tensor Cores (1,979 FP8 TFLOPS)
  • You need NVLink 4.0 for multi-GPU tensor parallelism
  • Cost-per-token: $1.89/hr รท 120 tok/s = $4.38/M tokens at batch=1 โ€” a latency-oriented figure; production batching cuts this substantially

Check live H100 rates โ€” observed spread is $1.89โ€“$3.49/hr across providers (~85%).

A100 SXM4: FP16/BF16 legacy

SpecValue
VRAM80 GB HBM2e (2.0 TB/s)
FP16 TFLOPS312
FP8 pathnone in silicon โ€” repo lists 624 as a derived/emulated figure, not a vendor spec
70B throughput~45 tok/s (site record: 70B INT8, batch=1, 2-GPU TP)
Observed price$1.59/hr (Lambda Labs)

The A100 is the right choice when:

  • You need BF16/FP16 (or INT8) โ€” not FP8: A100 has no FP8 tensor path, so "FP8 on A100" is emulation and should not be planned as such
  • You have existing A100 infrastructure and want to avoid migration
  • Your model is โ‰ค80 GB at FP16 (single GPU without TP)
  • 70B INT8 at TP=2: 2 ร— $1.59 = $3.18/hr รท 45 tok/s = $19.7/M at batch=1 โ€” the site's only A100 70B record

See A100-80GB specs for details.

L40S: Budget inference for smaller models

SpecValue
VRAM48 GB GDDR6
FP8 TFLOPS733
Throughput record~65 tok/s (8B FP8, batch=1) โ€” no 70B FP8 record exists
Observed price$0.69/hr (Vast.ai)

The L40S is the right choice when:

  • Your model fits in 48 GB: 7Bโ€“32B at FP8, 8B at FP16, 70B at INT4 (42.5 GB โ€” fits, 5.5 GB headroom)
  • Cost-per-token is the priority: 8B FP8 at $0.69/hr รท 65 tok/s = $2.95/M (batch=1)
  • You need a standard PCIe datacenter card (no NVLink โ€” multi-GPU is PCIe-bound)
  • 70B serving must stay single-GPU โ†’ INT4 is the only precision that fits; there is no published throughput record for that config, so measure it yourself

The old claim here โ€” "L40S delivers roughly the same 70B FP8 throughput as H100" โ€” was wrong and has been removed. A 70B FP8 model (81 GB) cannot run on an L40S (48 GB) at all; the L40S's only 70B-class record in the spec database is A100's 45 tok/s INT8 record, which belongs to a different GPU. See L40S specs.

When VRAM is the binding constraint

At 128K context, Llama 3.3 70B needs 87โ€“101 GB at FP8 (KV layer-count range above). None of H100, A100, or L40S fits it single-GPU. Your options:

  • H200 SXM5 (141 GB) โ€” single-GPU FP8 at 128K
  • 2ร— H100 with tensor parallelism โ€” NVLink overhead, ~2ร— hourly cost ($3.78/hr at Vast.ai rates)
  • INT4 quantization โ€” 42.5 GB service total: fits A100 (80), H100 (80), and L40S (48) single-GPU

The H100 vs H200 ROI analysis walks through the cost decision for 128K deployments.

Quantization changes the equation

INT4 compresses weights 4ร— relative to FP16 (weights-only figures; service totals add KV + ~7.5% overhead):

ModelFP16 (weights)INT4 (weights)Service INT4 totalFits on L40S (48GB)?
Llama 3.1 8B16 GB4 GB~5 GByes
Mistral 7B14 GB3.5 GB~4.5 GByes
Qwen 2.5 32B65 GB16 GB~18 GByes
Llama 3.3 70B141 GB35 GB42.5 GByes (5.5 GB headroom)
Qwen 2.5 72B145 GB36 GB~40 GByes

INT4 quality impact is workload-specific โ€” the quantization guide covers how to evaluate it.

Decision tree

  1. Model โ‰ค 30 GB service VRAM at your precision? โ†’ L40S (cheapest per-token: $2.95/M at 8B FP8 batch=1)
  2. 70B-class at FP8? โ†’ H100 (short context, single GPU) or 2ร— H100 TP (32K); H200 for 128K
  3. FP16/BF16 required and A100 already on hand? โ†’ A100 โ€” but do not plan "FP8 on A100"
  4. 70B on a budget, batch=1 acceptable? โ†’ L40S at INT4 (42.5 GB) โ€” benchmark your own throughput, no record exists
  5. VRAM exceeds 80 GB at FP8 (128K)? โ†’ H200 or multi-GPU H100

Limitations

  • Throughput figures are site GPU spec records (batch=1, named configs); combinations without records are marked as such rather than projected
  • VRAM figures exclude PagedAttention savings
  • GPU prices are [OBSERVED] as of 2026-10-05 and fluctuate
  • A100 has no native FP8 โ€” any FP8 number for it is derived, not measured
  • 70B KV cache is a 7โ€“20 GB range pending the repo's layer-count reconciliation

Related resources