Spec & Pricing Reference

NVIDIA L40S Cloud Pricing & Specs (2026)

The L40S offers 48 GB GDDR6 at 864 GB/s. Optimized for low-latency inference on 7B-32B models and multi-modal workloads where PCIe bandwidth is acceptable.

Target workload: Low-latency inference for 7B-32B models & multi-modal workloads

Memory48GB GDDR6
Bandwidth864 GB/s
FP8 TFLOPS733
Observed rate$0.00/hr
48 GB VRAM
Manufacturer SpecHIGH
SourceNVIDIA manufacturer specifications
VerifiedSep 30, 2026
Value
48 GB
Methodology

Manufacturer specification for onboard memory.

Refreshed daily from provider APIs and market scrapingSep 30, 2026
864 GB/s
Manufacturer SpecHIGH
SourceNVIDIA specification
VerifiedSep 30, 2026
Value
864 GB/s GB/s
Methodology

Peak memory bandwidth from the manufacturer specification.

Refreshed daily from provider APIs and market scrapingSep 30, 2026
$0.00/hr on-demand
Observed On-Demand RateHIGH
SourceObserved provider API rate (lowest on-demand row)
VerifiedSep 30, 2026
Value
0 USD/hr
Methodology

Lowest on-demand hourly row for "L40S" across tracked providers in data/providers.json; refreshed daily.

Refreshed daily from provider APIs and market scrapingSep 30, 2026
Methodology β†’

NVIDIA L40S: Key Numbers at a Glance

Rental cost: lowest observed on-demand $0.00/hr, spot rows from $0.00/hr across 4 tracked providers (refreshed daily).

70B fit: Fits Llama 3.3 70B at INT4 (min 40 GB including KV-cache) at short context only; FP8/FP16 do not fit single-GPU.

Bottleneck: Severely memory-bandwidth bound β€” GDDR6 bus cannot feed FP8 Tensor Core demand at batchβ‰₯16

Observed Pricing for NVIDIA L40S

Spot and on-demand rows for L40S refreshed from provider APIs. Lowest on-demand: $0.00/hr.

ProviderGPU & VRAMInterconnectSpot Price ($/hr)On-Demand ($/hr)Monthly ($/720h)StatusAction
Lambda Labs
N/A$0/hr
Calculated EstimateMEDIUM
SourceLambda Labs
VerifiedSep 26, 2026
Value
0 /hr
Methodology

Observed spot rate from public cloud APIs. On-demand calculated at 2.5Γ— spot (conservative; actual range 1.8×–3.2Γ—).

Assumptions & Parameters
  • range: 1.8×–3.2Γ—
  • note: Spot rates fluctuate
Refreshed daily from provider APIs and market scrapingSep 26, 2026
$0 / hr$0 / moInstant
Vast.ai
N/A$0.45/hr
Calculated EstimateMEDIUM
SourceVast.ai
VerifiedSep 26, 2026
Value
0.45 /hr
Methodology

Observed spot rate from public cloud APIs. On-demand calculated at 2.5Γ— spot (conservative; actual range 1.8×–3.2Γ—).

Assumptions & Parameters
  • range: 1.8×–3.2Γ—
  • note: Spot rates fluctuate
Refreshed daily from provider APIs and market scrapingSep 26, 2026
$0.65 / hr$0 / moInstant
RunPod
N/A$0.86/hr
Calculated EstimateMEDIUM
SourceRunPod
VerifiedSep 26, 2026
Value
0.86 /hr
Methodology

Observed spot rate from public cloud APIs. On-demand calculated at 2.5Γ— spot (conservative; actual range 1.8×–3.2Γ—).

Assumptions & Parameters
  • range: 1.8×–3.2Γ—
  • note: Spot rates fluctuate
Refreshed daily from provider APIs and market scrapingSep 26, 2026
$0.86 / hr$0 / moInstant
Spheron
N/A$1.07/hr
Calculated EstimateMEDIUM
SourceSpheron
VerifiedSep 26, 2026
Value
1.07 /hr
Methodology

Observed spot rate from public cloud APIs. On-demand calculated at 2.5Γ— spot (conservative; actual range 1.8×–3.2Γ—).

Assumptions & Parameters
  • range: 1.8×–3.2Γ—
  • note: Spot rates fluctuate
Refreshed daily from provider APIs and market scrapingSep 26, 2026
$0.96 / hr$0 / moInstant
Monthly estimates use continuous spot rate Γ— 720h. On-demand calculated at 2.5Γ— spot (range: 1.8×–3.2Γ—). All rates subject to preemption and provider availability.
Data Freshness: Public Cloud APIs & Market Scraping | Refreshed Daily (UTC)Benchmark Baseline: Ubuntu 24.04, CUDA 12.4, vLLM v0.6.x, PagedAttention v2, FlashAttention-3

Compatible Models for NVIDIA L40S

Models from the VRAM registry whose minimum INT4 footprint (weights + KV-cache + runtime overhead) fits 48 GB. FP16 shows where full precision also fits single-GPU.

Browse all model VRAM pages β†’

Specifications

ArchitectureAda Lovelace AD102 β€” 5nm TSMC
Memory48GB GDDR6
Bandwidth864 GB/s
InterconnectPCIe 4.0 (64 GB/s)
TDP350W
FP16 TFLOPS366
FP8 TFLOPS733
FP4 TFLOPSN/A
Recommended quantizationAWQ / GPTQ / GGUF β€” INT4 essential for 30B+ models
Best cluster topologyPCIe Single Node β€” no NVLink, 2-4 GPU max

L40S Ada Lovelace FP8 Inference vs Training Limitations: The L40S excels at FP8 inference β€” Ada Lovelace's native FP8 Tensor Core support delivers 733 TFLOPS, making it the most cost-effective FP8 inference GPU below the H100 class. For Llama 8B FP8, the L40S achieves ~65 tok/s at batch=1 with 48 GB VRAM leaving 44 GB for KV-cache. However, the L40S has critical limitations for training: (1) No FP64 support β€” scientific computing and double-precision gradient accumulation are impossible. (2) No NVLink β€” multi-GPU training relies on PCIe 4.0 (64 GB/s), limiting tensor parallelism to 2 GPUs as communication overhead becomes significant. (3) 864 GB/s GDDR6 bandwidth is 4x slower than H100's 3.35 TB/s HBM3 β€” batch inference scales poorly beyond batch=16. The optimal use case: production inference for 7B-30B models where FP8 precision provides the best tokens/dollar ratio. Not suitable for training or fine-tuning.

Next steps

Compare Alternatives

Head-to-head comparisons against this GPU β€” specs, observed hourly rates, and the workload verdict.

Related GPUs

Frequently Asked Questions

How much does it cost to rent NVIDIA L40S per hour?β–Ύ
The lowest observed on-demand rate for NVIDIA L40S is $0.00/hr (4 tracked providers); observed spot rows start at $0.00/hr. Rates refresh daily from provider APIs.
Can NVIDIA L40S run 70B parameter LLMs?β–Ύ
Fits Llama 3.3 70B at INT4 (min 40 GB including KV-cache) at short context only; FP8/FP16 do not fit single-GPU.
What is NVIDIA L40S's memory bandwidth?β–Ύ
NVIDIA L40S has 864 GB/s of memory bandwidth across 48GB GDDR6. Token generation is memory-bandwidth bound at batch size 1, so decode throughput scales with this figure β€” no tokens/sec value is claimed without a benchmark row.
Rates observed across 4 providers; refreshed daily (UTC).Spec sources: NVIDIA manufacturer specifications.Methodology β†’