NVIDIA A100 40GB Cloud Pricing & Specs (2026)
The A100 40GB pairs Ampere tensor cores with 40 GB of HBM2 at 1.55 TB/s and NVLink 3.0 at 600 GB/s. It runs 13B-30B class models at INT8/INT4 with room for KV-cache, at a lower price point than the 80GB variant in fleets that still split the SKU.
Target workload: Budget 13B-30B inference & LoRA fine-tuning
40 GB VRAM
Manufacturer specifications. OpenGPU Radar does not currently track A100 40GB rate rows — no price is asserted.
1.55 TB/s
Peak memory bandwidth from the manufacturer specification.
NVIDIA A100 40GB: Key Numbers at a Glance
Rental cost: no verified rate rows are tracked yet — this page asserts specifications only, never a price.
70B fit: Fits Llama 3.3 70B at INT4 (min 40 GB including KV-cache) at short context only; FP8/FP16 do not fit single-GPU.
Bottleneck: Memory-bandwidth bound — 1.55 TB/s caps batch scaling before compute saturates
OpenGPU Radar does not currently track this GPU rows in data/providers.json. No price is asserted on this page — the specifications below are manufacturer-sourced.
Compare GPUs with tracked rates →Compatible Models for NVIDIA A100 40GB
Models from the VRAM registry whose minimum INT4 footprint (weights + KV-cache + runtime overhead) fits 40 GB. FP16 shows where full precision also fits single-GPU.
VRAM breakdown →
VRAM breakdown →
VRAM breakdown →
VRAM breakdown →
VRAM breakdown →
VRAM breakdown →
VRAM breakdown →
VRAM breakdown →
Specifications
The A100 40GB is an Ampere part: it has no FP8 tensor cores, so FP8-native models must run at FP16/INT8 (or INT4 post-training quantization). 40 GB fits 30B-class INT4 weights (≈16-18 GB) with substantial KV-cache headroom, and 13B FP16 comfortably. NVLink 3.0 at 600 GB/s supports 2-4 way tensor parallelism when a workload exceeds 40 GB, but 40 GB caps Llama 70B (38 GB INT4) with no context budget — for 70B use the 80GB variant or H100/H200. Memory bandwidth of 1.55 TB/s is the practical throughput ceiling: token generation is bandwidth-bound at any batch size.
Next steps