Spec & Pricing Reference

NVIDIA T4 Cloud Pricing & Specs (2026)

The T4 is Turing's 70W inference card with 16 GB GDDR6 at 0.32 TB/s. It still backs a long tail of inference, embedding, and reranker deployments where power and cost per instance matter more than decode speed.

Target workload: Edge inference for ≤7B models, embeddings & rerankers

Memory16GB GDDR6
Bandwidth0.32 TB/s
FP8 TFLOPSNot supported
Observed rateNo tracked rows
16 GB VRAM
Manufacturer SpecHIGH
SourceNVIDIA Tesla T4 datasheet
VerifiedSep 30, 2026
Value
16 GB
Methodology

Manufacturer specifications. OpenGPU Radar does not currently track T4 rate rows — no price is asserted.

Refreshed daily from provider APIs and market scrapingSep 30, 2026
0.32 TB/s
Manufacturer SpecHIGH
SourceNVIDIA specification
VerifiedSep 30, 2026
Value
0.32 TB/s GB/s
Methodology

Peak memory bandwidth from the manufacturer specification.

Refreshed daily from provider APIs and market scrapingSep 30, 2026
Methodology →

NVIDIA T4: Key Numbers at a Glance

Rental cost: no verified rate rows are tracked yet — this page asserts specifications only, never a price.

70B fit: Does not fit Llama 3.3 70B single-GPU — INT4 alone needs 40 GB vs 16 GB available. 7B-13B class is the single-GPU target.

Bottleneck: Memory-bandwidth bound — 0.32 TB/s caps interactive generation

Specs Only — No Rate Rows Tracked
Pricing coverage for NVIDIA T4

OpenGPU Radar does not currently track this GPU rows in data/providers.json. No price is asserted on this page — the specifications below are manufacturer-sourced.

Compare GPUs with tracked rates →

Compatible Models for NVIDIA T4

Models from the VRAM registry whose minimum INT4 footprint (weights + KV-cache + runtime overhead) fits 16 GB. FP16 shows where full precision also fits single-GPU.

Browse all model VRAM pages →

Specifications

ArchitectureTuring — 12nm TSMC
Memory16GB GDDR6
Bandwidth0.32 TB/s
InterconnectPCIe 3.0 (16 GB/s)
TDP70W
FP16 TFLOPS65
FP8 TFLOPSN/A — no FP8 Tensor Cores (Turing generation)
FP4 TFLOPSN/A
Recommended quantizationGGUF / AWQ — INT4 for 7B class
Best cluster topologyPCIe single node — edge inference appliance

The T4 is a Turing card: FP16 tensor cores only (no FP8, no FP4), 16 GB of VRAM, and 0.32 TB/s of memory bandwidth. It fits 7B INT4 models (≈4 GB weights) with KV-cache room for chat-length contexts, and is well matched to embedding and reranker models that are batch-processed rather than decode-bound. 13B INT4 (≈7 GB) fits but with reduced context headroom; anything 30B+ does not fit. At 0.32 TB/s, interactive token generation on 7B models is the practical ceiling — throughput-class workloads belong on 0.6+ TB/s cards. No rate rows are tracked for the T4 in OpenGPU Radar's provider table.

Next steps

Related GPUs

Frequently Asked Questions

How much does it cost to rent NVIDIA T4 per hour?▾
OpenGPU Radar does not currently track rate rows for NVIDIA T4. The specifications on this page are manufacturer-sourced; no price is asserted.
Can NVIDIA T4 run 70B parameter LLMs?▾
Does not fit Llama 3.3 70B single-GPU — INT4 alone needs 40 GB vs 16 GB available. 7B-13B class is the single-GPU target.
What is NVIDIA T4's memory bandwidth?▾
NVIDIA T4 has 0.32 TB/s of memory bandwidth across 16GB GDDR6. Token generation is memory-bandwidth bound at batch size 1, so decode throughput scales with this figure — no tokens/sec value is claimed without a benchmark row.
Specifications only — no rate rows tracked for this GPU.Spec sources: NVIDIA Tesla T4 datasheet.Methodology →