Spec & Pricing Reference

NVIDIA GeForce RTX 4090 Cloud Pricing & Specs (2026)

The RTX 4090 delivers 82.6 TFLOPS FP16 at the lowest spot pricing. Ideal for QLoRA fine-tuning, Stable Diffusion, and development workloads where PCIe bandwidth is sufficient.

Target workload: Budget fine-tuning, QLoRA, Stable Diffusion, & local dev

Memory24GB GDDR6X
Bandwidth1.0 TB/s
FP8 TFLOPS165
Observed rateNo tracked rows
24 GB VRAM
Manufacturer SpecHIGH
SourceNVIDIA manufacturer specifications
VerifiedSep 30, 2026
Value
24 GB
Methodology

Manufacturer specification for onboard memory.

Refreshed daily from provider APIs and market scrapingSep 30, 2026
1.0 TB/s
Manufacturer SpecHIGH
SourceNVIDIA specification
VerifiedSep 30, 2026
Value
1.0 TB/s GB/s
Methodology

Peak memory bandwidth from the manufacturer specification.

Refreshed daily from provider APIs and market scrapingSep 30, 2026
Methodology →

NVIDIA GeForce RTX 4090: Key Numbers at a Glance

Rental cost: no verified rate rows are tracked yet — this page asserts specifications only, never a price.

70B fit: Does not fit Llama 3.3 70B single-GPU — INT4 alone needs 40 GB vs 24 GB available. 7B-13B class is the single-GPU target.

Bottleneck: Extremely memory-bandwidth bound — GDDR6X 1 TB/s cannot sustain FP16 inference at batch≥4

Specs Only — No Rate Rows Tracked
Pricing coverage for NVIDIA GeForce RTX 4090

OpenGPU Radar does not currently track "RTX 4090" rows in data/providers.json. No price is asserted on this page — the specifications below are manufacturer-sourced.

Compare GPUs with tracked rates →

Compatible Models for NVIDIA GeForce RTX 4090

Models from the VRAM registry whose minimum INT4 footprint (weights + KV-cache + runtime overhead) fits 24 GB. FP16 shows where full precision also fits single-GPU.

Browse all model VRAM pages →

Specifications

ArchitectureAda Lovelace AD102 — 5nm TSMC
Memory24GB GDDR6X
Bandwidth1.0 TB/s
InterconnectPCIe 4.0 (64 GB/s)
TDP450W
FP16 TFLOPS82.6
FP8 TFLOPS165
FP4 TFLOPSN/A
Recommended quantizationGGUF / AWQ / GPTQ — INT4 mandatory for 13B+
Best cluster topologyPCIe Single Node — 2 GPU max for training

RTX 4090 PCIe Gen4 Lane Saturation & QLoRA Analysis: The RTX 4090 connects via PCIe 4.0 x16 (64 GB/s bidirectional) — a critical bottleneck for multi-GPU workloads. Unlike SXM GPUs with NVLink, the 4090 lacks P2P NVLink: 2x RTX 4090 training requires PCIe round-trips through the CPU, yielding only 1.6x speedup (not 2x). For QLoRA fine-tuning, this is acceptable: the 24 GB GDDR6X fits Llama 70B INT4 (~14 GB) with 10 GB remaining for adapter weights and optimizer states in CPU RAM. The 1.0 TB/s memory bandwidth supports ~28 tok/s at INT4 — sufficient for development iteration. Host stability risks: consumer-grade GPUs on cloud spot markets may have varying thermal conditions, driver versions, and PCIe slot configurations. We recommend validating GPU health via nvidia-smi before deploying production workloads. The sweet spot: QLoRA fine-tuning of 13B-30B models where the 24 GB VRAM is adequate and the $0.34-$0.69/hr spot pricing beats datacenter GPUs by 5-10x.

Next steps

Compare Alternatives

Head-to-head comparisons against this GPU — specs, observed hourly rates, and the workload verdict.

Related GPUs

Frequently Asked Questions

How much does it cost to rent NVIDIA GeForce RTX 4090 per hour?▾
OpenGPU Radar does not currently track rate rows for NVIDIA GeForce RTX 4090. The specifications on this page are manufacturer-sourced; no price is asserted.
Can NVIDIA GeForce RTX 4090 run 70B parameter LLMs?▾
Does not fit Llama 3.3 70B single-GPU — INT4 alone needs 40 GB vs 24 GB available. 7B-13B class is the single-GPU target.
What is NVIDIA GeForce RTX 4090's memory bandwidth?▾
NVIDIA GeForce RTX 4090 has 1.0 TB/s of memory bandwidth across 24GB GDDR6X. Token generation is memory-bandwidth bound at batch size 1, so decode throughput scales with this figure — no tokens/sec value is claimed without a benchmark row.
Specifications only — no rate rows tracked for this GPU.Spec sources: NVIDIA manufacturer specifications.Methodology →