NVIDIA A100 80GB SXM4 Cloud Pricing & Specs (2026)
The A100 80GB remains the cost-effective choice for mid-size LLM fine-tuning and computer vision. Its 80 GB HBM2e and 2 TB/s bandwidth handle 13B-30B parameter models efficiently.
Target workload: Cost-effective mid-size LLM fine-tuning & computer vision
80 GB VRAM
Manufacturer specification for onboard memory.
2.0 TB/s
Peak memory bandwidth from the manufacturer specification.
NVIDIA A100 80GB SXM4: Key Numbers at a Glance
Rental cost: no verified rate rows are tracked yet — this page asserts specifications only, never a price.
70B fit: Fits Llama 3.3 70B at FP8/INT8 (weights ≈71 GB at 1 byte/param) with a short-context KV budget; FP16 (140 GB) requires 2-way tensor parallelism.
Bottleneck: Memory-bandwidth bound — 2.0 TB/s limits batch scaling
OpenGPU Radar does not currently track "A100" rows in data/providers.json. No price is asserted on this page — the specifications below are manufacturer-sourced.
Compare GPUs with tracked rates →Compatible Models for NVIDIA A100 80GB SXM4
Models from the VRAM registry whose minimum INT4 footprint (weights + KV-cache + runtime overhead) fits 80 GB. FP16 shows where full precision also fits single-GPU.
VRAM breakdown →
VRAM breakdown →
VRAM breakdown →
VRAM breakdown →
VRAM breakdown →
VRAM breakdown →
VRAM breakdown →
VRAM breakdown →
Specifications
The A100 is memory-bandwidth bound across all batch sizes. Its 2.0 TB/s HBM2e provides only ~60% of the bandwidth needed to sustain FP16 tensor core utilization at batch≥32. For Llama 70B, the 80 GB VRAM fits the model in FP16 with ~8 GB remaining for KV-cache (limited to 8k context at batch=1). Tensor parallelism across 2-4 GPUs is mandatory for 70B models. The A100's strength is cost efficiency: at $1.59/hr on Lambda Labs, it delivers the best $/token for 13B-30B models.
Next steps
Compare Alternatives
Head-to-head comparisons against this GPU — specs, observed hourly rates, and the workload verdict.
H100 SXM5 leads FP8-native serving: 1,979 FP8 TFLOPS and 3.35 TB/s bandwidth deliver a higher decode ceiling per GPU, and NVLink 4.0 at 900 GB/s scales tensor parallelism. The A100 80GB is 16% cheaper per hour observed ($1.59 vs $1.89) and remains the budget choice for INT8/INT4 batch workloads and Ampere-standardized tooling — but it has no FP8 tensor cores, so Hopper's native FP8 pipeline advantage does not exist on Ampere.
Different classes: the RTX 4090 is 4.7× cheaper observed ($0.34 vs $1.59/hr) and runs 13B-class INT4 models with KV headroom — ideal for prototyping, small-model LoRA, and low-traffic serving. The A100's 80 GB fits 70B INT8/INT4 workloads and 4-way NVLink tensor parallelism that the 4090 cannot touch (24 GB, PCIe-only). Buy the 4090 while the model fits 24 GB; buy the A100 when it doesn't.
L40S is 57% cheaper observed ($0.69 vs $1.59/hr) and FP8-capable at 733 TFLOPS — the value pick for ≤30B FP8/INT4 serving and LoRA fine-tuning. The A100's 80 GB versus the L40S's 48 GB is the dividing line: 70B-class stacks exceed 48 GB with context, and the A100's NVLink 3.0 enables 2-4 way tensor parallelism where the L40S is PCIe-only. Choose L40S for small/medium models; A100 for 70B or multi-GPU.