Spec & Pricing Reference

NVIDIA A100 40GB Cloud Pricing & Specs (2026)

The A100 40GB pairs Ampere tensor cores with 40 GB of HBM2 at 1.55 TB/s and NVLink 3.0 at 600 GB/s. It runs 13B-30B class models at INT8/INT4 with room for KV-cache, at a lower price point than the 80GB variant in fleets that still split the SKU.

Target workload: Budget 13B-30B inference & LoRA fine-tuning

Memory40GB HBM2
Bandwidth1.55 TB/s
FP8 TFLOPSNot supported
Observed rateNo tracked rows
40 GB VRAM
Manufacturer SpecHIGH
SourceNVIDIA A100 40GB datasheet
VerifiedSep 30, 2026
Value
40 GB
Methodology

Manufacturer specifications. OpenGPU Radar does not currently track A100 40GB rate rows — no price is asserted.

Refreshed daily from provider APIs and market scrapingSep 30, 2026
1.55 TB/s
Manufacturer SpecHIGH
SourceNVIDIA specification
VerifiedSep 30, 2026
Value
1.55 TB/s GB/s
Methodology

Peak memory bandwidth from the manufacturer specification.

Refreshed daily from provider APIs and market scrapingSep 30, 2026
Methodology →

NVIDIA A100 40GB: Key Numbers at a Glance

Rental cost: no verified rate rows are tracked yet — this page asserts specifications only, never a price.

70B fit: Fits Llama 3.3 70B at INT4 (min 40 GB including KV-cache) at short context only; FP8/FP16 do not fit single-GPU.

Bottleneck: Memory-bandwidth bound — 1.55 TB/s caps batch scaling before compute saturates

Specs Only — No Rate Rows Tracked
Pricing coverage for NVIDIA A100 40GB

OpenGPU Radar does not currently track this GPU rows in data/providers.json. No price is asserted on this page — the specifications below are manufacturer-sourced.

Compare GPUs with tracked rates →

Compatible Models for NVIDIA A100 40GB

Models from the VRAM registry whose minimum INT4 footprint (weights + KV-cache + runtime overhead) fits 40 GB. FP16 shows where full precision also fits single-GPU.

Browse all model VRAM pages →

Specifications

ArchitectureAmpere GA100 — 7nm TSMC
Memory40GB HBM2
Bandwidth1.55 TB/s
InterconnectNVLink 3.0 (600 GB/s)
TDP400W
FP16 TFLOPS312
FP8 TFLOPSN/A — no FP8 Tensor Cores (Ampere generation)
FP4 TFLOPSN/A
Recommended quantizationINT8 / INT4 (GPTQ/AWQ) — no FP8 support
Best cluster topologyNVLink 3.0 baseboard (TP=2/4) or PCIe single node

The A100 40GB is an Ampere part: it has no FP8 tensor cores, so FP8-native models must run at FP16/INT8 (or INT4 post-training quantization). 40 GB fits 30B-class INT4 weights (≈16-18 GB) with substantial KV-cache headroom, and 13B FP16 comfortably. NVLink 3.0 at 600 GB/s supports 2-4 way tensor parallelism when a workload exceeds 40 GB, but 40 GB caps Llama 70B (38 GB INT4) with no context budget — for 70B use the 80GB variant or H100/H200. Memory bandwidth of 1.55 TB/s is the practical throughput ceiling: token generation is bandwidth-bound at any batch size.

Next steps

Related GPUs

Frequently Asked Questions

How much does it cost to rent NVIDIA A100 40GB per hour?▾
OpenGPU Radar does not currently track rate rows for NVIDIA A100 40GB. The specifications on this page are manufacturer-sourced; no price is asserted.
Can NVIDIA A100 40GB run 70B parameter LLMs?▾
Fits Llama 3.3 70B at INT4 (min 40 GB including KV-cache) at short context only; FP8/FP16 do not fit single-GPU.
What is NVIDIA A100 40GB's memory bandwidth?▾
NVIDIA A100 40GB has 1.55 TB/s of memory bandwidth across 40GB HBM2. Token generation is memory-bandwidth bound at batch size 1, so decode throughput scales with this figure — no tokens/sec value is claimed without a benchmark row.
Specifications only — no rate rows tracked for this GPU.Spec sources: NVIDIA A100 40GB datasheet.Methodology →