NVIDIA L40S Cloud Pricing & Specs (2026)
The L40S offers 48 GB GDDR6 at 864 GB/s. Optimized for low-latency inference on 7B-32B models and multi-modal workloads where PCIe bandwidth is acceptable.
Target workload: Low-latency inference for 7B-32B models & multi-modal workloads
48 GB VRAM
Manufacturer specification for onboard memory.
864 GB/s
Peak memory bandwidth from the manufacturer specification.
$0.00/hr on-demand
Lowest on-demand hourly row for "L40S" across tracked providers in data/providers.json; refreshed daily.
NVIDIA L40S: Key Numbers at a Glance
Rental cost: lowest observed on-demand $0.00/hr, spot rows from $0.00/hr across 4 tracked providers (refreshed daily).
70B fit: Fits Llama 3.3 70B at INT4 (min 40 GB including KV-cache) at short context only; FP8/FP16 do not fit single-GPU.
Bottleneck: Severely memory-bandwidth bound β GDDR6 bus cannot feed FP8 Tensor Core demand at batchβ₯16
Observed Pricing for NVIDIA L40S
Spot and on-demand rows for L40S refreshed from provider APIs. Lowest on-demand: $0.00/hr.
Featured GPU Pricing Pages
80GB HBM3 β’ 3.35 TB/s β’ NVLink 4.0
40GB HBM2 β’ 1.55 TB/s β’ NVLink 3.0
Verified spot rates & reserved pricing
24GB GDDR6 β’ 0.62 TB/s β’ PCIe 4.0
24GB GDDR6X β’ 0.94 TB/s β’ PCIe 4.0
16GB GDDR6 β’ 0.32 TB/s β’ PCIe 3.0
Compatible Models for NVIDIA L40S
Models from the VRAM registry whose minimum INT4 footprint (weights + KV-cache + runtime overhead) fits 48 GB. FP16 shows where full precision also fits single-GPU.
VRAM breakdown β
VRAM breakdown β
VRAM breakdown β
VRAM breakdown β
VRAM breakdown β
VRAM breakdown β
VRAM breakdown β
VRAM breakdown β
Specifications
L40S Ada Lovelace FP8 Inference vs Training Limitations: The L40S excels at FP8 inference β Ada Lovelace's native FP8 Tensor Core support delivers 733 TFLOPS, making it the most cost-effective FP8 inference GPU below the H100 class. For Llama 8B FP8, the L40S achieves ~65 tok/s at batch=1 with 48 GB VRAM leaving 44 GB for KV-cache. However, the L40S has critical limitations for training: (1) No FP64 support β scientific computing and double-precision gradient accumulation are impossible. (2) No NVLink β multi-GPU training relies on PCIe 4.0 (64 GB/s), limiting tensor parallelism to 2 GPUs as communication overhead becomes significant. (3) 864 GB/s GDDR6 bandwidth is 4x slower than H100's 3.35 TB/s HBM3 β batch inference scales poorly beyond batch=16. The optimal use case: production inference for 7B-30B models where FP8 precision provides the best tokens/dollar ratio. Not suitable for training or fine-tuning.
Next steps
Compare Alternatives
Head-to-head comparisons against this GPU β specs, observed hourly rates, and the workload verdict.
The RTX 4090 offers lower hourly cost for INT4 inference. The L40S leads on FP8 precision, memory capacity, enterprise reliability, and ECC memory for production workloads.
L40S is 57% cheaper observed ($0.69 vs $1.59/hr) and FP8-capable at 733 TFLOPS β the value pick for β€30B FP8/INT4 serving and LoRA fine-tuning. The A100's 80 GB versus the L40S's 48 GB is the dividing line: 70B-class stacks exceed 48 GB with context, and the A100's NVLink 3.0 enables 2-4 way tensor parallelism where the L40S is PCIe-only. Choose L40S for small/medium models; A100 for 70B or multi-GPU.