NVIDIA T4 Cloud Pricing & Specs (2026)
The T4 is Turing's 70W inference card with 16 GB GDDR6 at 0.32 TB/s. It still backs a long tail of inference, embedding, and reranker deployments where power and cost per instance matter more than decode speed.
Target workload: Edge inference for ≤7B models, embeddings & rerankers
16 GB VRAM
Manufacturer specifications. OpenGPU Radar does not currently track T4 rate rows — no price is asserted.
0.32 TB/s
Peak memory bandwidth from the manufacturer specification.
NVIDIA T4: Key Numbers at a Glance
Rental cost: no verified rate rows are tracked yet — this page asserts specifications only, never a price.
70B fit: Does not fit Llama 3.3 70B single-GPU — INT4 alone needs 40 GB vs 16 GB available. 7B-13B class is the single-GPU target.
Bottleneck: Memory-bandwidth bound — 0.32 TB/s caps interactive generation
OpenGPU Radar does not currently track this GPU rows in data/providers.json. No price is asserted on this page — the specifications below are manufacturer-sourced.
Compare GPUs with tracked rates →Compatible Models for NVIDIA T4
Models from the VRAM registry whose minimum INT4 footprint (weights + KV-cache + runtime overhead) fits 16 GB. FP16 shows where full precision also fits single-GPU.
VRAM breakdown →
VRAM breakdown →
VRAM breakdown →
VRAM breakdown →
Specifications
The T4 is a Turing card: FP16 tensor cores only (no FP8, no FP4), 16 GB of VRAM, and 0.32 TB/s of memory bandwidth. It fits 7B INT4 models (≈4 GB weights) with KV-cache room for chat-length contexts, and is well matched to embedding and reranker models that are batch-processed rather than decode-bound. 13B INT4 (≈7 GB) fits but with reduced context headroom; anything 30B+ does not fit. At 0.32 TB/s, interactive token generation on 7B models is the practical ceiling — throughput-class workloads belong on 0.6+ TB/s cards. No rate rows are tracked for the T4 in OpenGPU Radar's provider table.
Next steps