NVIDIA L40S Cloud Pricing & Rental Rates
The L40S offers 48 GB GDDR6 at 864 GB/s. Optimized for low-latency inference on 7B-32B models and multi-modal workloads where PCIe bandwidth is acceptable.
Target Workload: Low-latency inference for 7B-32B models & multi-modal workloads
Live Cloud Pricing for NVIDIA L40S
Spot and reserved rates refreshed from provider APIs. Click a provider to deploy.
| Provider | GPU & VRAM | Interconnect | Spot Rate | On-Demand | Monthly | Status | Action | |
|---|---|---|---|---|---|---|---|---|
Community | PCIe 4.0 (64 GB/s) | $0.69 / hr | $1.73 / hr | $422 / mo | Instant | |||
Cloud | PCIe 4.0 (64 GB/s) | $1.09 / hr | $2.73 / hr | $667 / mo | Instant | |||
Bare Metal | PCIe 4.0 (64 GB/s) | $1.19 / hr | $2.97 / hr | $728 / mo | Instant | |||
Dedicated | PCIe 4.0 (64 GB/s) | $1.49 / hr | $3.73 / hr | $912 / mo | Instant |
Models That Fit NVIDIA L40S (48GB GDDR6)
Deterministic VRAM calculation (FP8) via entity graph. Only editorial and enriched models shown.
| MODEL | PARAMS | CONTEXT | FP8 VRAM | FIT? | CALCULATOR |
|---|---|---|---|---|---|
| DeepSeek R1 Distill Qwen 32B | 32B | 128K | 42.9 GB | โ fits | Pre-filled โ |
| Llama 3.1 8B Instruct | 8.03B | 128K | 10.8 GB | โ fits | Pre-filled โ |
| Llama 3.2 3B Instruct | 3.2B | 128K | 4.3 GB | โ fits | Pre-filled โ |
| Llama 3.2 1B Instruct | 1B | 128K | 1.3 GB | โ fits | Pre-filled โ |
| Qwen 2.5 Coder 32B | 32.5B | 128K | 43.6 GB | โ fits | Pre-filled โ |
| Qwen 2.5 Coder 14B | 14.7B | 128K | 19.7 GB | โ fits | Pre-filled โ |
| Qwen 2.5 Coder 7B | 7.61B | 128K | 10.2 GB | โ fits | Pre-filled โ |
| Qwen 2.5 14B Instruct | 14.7B | 128K | 19.7 GB | โ fits | Pre-filled โ |
| Qwen 2.5 7B Instruct | 7.61B | 128K | 10.2 GB | โ fits | Pre-filled โ |
| Mistral NeMo 12B | 12B | 128K | 16.1 GB | โ fits | Pre-filled โ |
| Gemma 2 27B | 27B | 8K | 35.7 GB | โ fits | Pre-filled โ |
| Gemma 2 9B | 9.24B | 8K | 12.2 GB | โ fits | Pre-filled โ |
Inference & Serving Capacity
Practical model feasibility, max batch sizes, and KV-cache retention limits for NVIDIA L40S.
Llama 3.3 70B
OOM48 GB insufficient for 70B at any precision
DeepSeek 671B
OOMRequires multi-GPU cluster
Qwen 2.5 32B
FEASIBLEAWQ/GPTQ mandatory, 30B max practical
vLLM Throughput (FP8)
~65 tok/s (vLLM, Llama 8B FP8, batch=1)
Estimated tokens/second, single GPU, Llama-class model
Max Context Window (Llama 70B)
16k tokens (INT4) โ requires quantization, tight KV-cache budget
Maximum context length before KV-cache eviction
Hardware Bottleneck Analysis
Whether NVIDIA L40S is compute-bound (TFLOPS) or memory-bandwidth bound (GB/s) across workloads.
Bottleneck Classification
Severely memory-bandwidth bound โ 864 GB/s vs 733 TFLOPS FP8
Recommended Quantization
AWQ / GPTQ / GGUF โ INT4 essential for 30B+ models
Best Cluster Topology
PCIe Single Node โ no NVLink, 2-4 GPU max
Deep Analysis
L40S Ada Lovelace FP8 Inference vs Training Limitations: The L40S excels at FP8 inference โ Ada Lovelace's native FP8 Tensor Core support delivers 733 TFLOPS, making it the most cost-effective FP8 inference GPU below the H100 class. For Llama 8B FP8, the L40S achieves ~65 tok/s at batch=1 with 48 GB VRAM leaving 44 GB for KV-cache. However, the L40S has critical limitations for training: (1) No FP64 support โ scientific computing and double-precision gradient accumulation are impossible. (2) No NVLink โ multi-GPU training relies on PCIe 4.0 (64 GB/s), limiting tensor parallelism to 2 GPUs as communication overhead becomes significant. (3) 864 GB/s GDDR6 bandwidth is 4x slower than H100's 3.35 TB/s HBM3 โ batch inference scales poorly beyond batch=16. The optimal use case: production inference for 7B-30B models where FP8 precision provides the best tokens/dollar ratio. Not suitable for training or fine-tuning.
Architecture & Die Breakdown
Architecture
Ada Lovelace AD102 โ 5nm TSMC
TDP
350W
Memory Subsystem
48GB GDDR6 at 864 GB/s bandwidth. GDDR6/X provides cost-effective bandwidth for workloads that don't require HBM-level throughput.
Interconnect
PCIe 4.0 (64 GB/s). Standard PCIe bus. Suitable for single-GPU workloads or multi-GPU training with gradient accumulation.
Precision Performance
| Precision | TFLOPS | Use Case |
|---|---|---|
| FP8 | 733 | Training & inference with mixed-precision |
| FP16 | 366 | Full-precision training, fine-tuning, evaluation |
Break-Even ROI Calculator
Monthly hours where reserved pricing beats spot for NVIDIA L40S. Above the break-even point, reserve commits save money.
Spheron
RunPod
Lambda Labs
Break-even at 200 hours/month: if you run NVIDIA L40S more than 200 hours per month, reserved pricing on all three providers saves money. At 720 hours/month (24/7), reserved saves $187+/mo vs on-demand.
Related GPUs
Frequently Asked Questions
How much does it cost to rent an NVIDIA L40S per hour?โพ
What is the monthly reserved pricing for NVIDIA L40S?โพ
Can an NVIDIA L40S run 70B parameter LLMs?โพ
Is spot pricing reliable for distributed NVIDIA L40S training?โพ
Compare alternatives & next steps