NVIDIA A10G Cloud Pricing & Specs (2026)
The A10G is Ampere's cloud-inference SKU: 24 GB GDDR6 at 0.62 TB/s in a 150W envelope. It is the GPU behind AWS g5 instances and similar inference-optimized fleets, tuned for 7B-13B INT4 serving rather than training.
Target workload: Cloud inference for 7B-13B models & graphics rendering
24 GB VRAM
Spec fields mirrored from the verified telemetry row; board attributes per NVIDIA A10G datasheet. No rate rows tracked.
0.62 TB/s
Peak memory bandwidth from the manufacturer specification.
NVIDIA A10G: Key Numbers at a Glance
Rental cost: no verified rate rows are tracked yet — this page asserts specifications only, never a price.
70B fit: Does not fit Llama 3.3 70B single-GPU — INT4 alone needs 40 GB vs 24 GB available. 7B-13B class is the single-GPU target.
Bottleneck: Memory-bandwidth bound — 0.62 TB/s is the token-generation ceiling
OpenGPU Radar does not currently track this GPU rows in data/providers.json. No price is asserted on this page — the specifications below are manufacturer-sourced.
Compare GPUs with tracked rates →Compatible Models for NVIDIA A10G
Models from the VRAM registry whose minimum INT4 footprint (weights + KV-cache + runtime overhead) fits 24 GB. FP16 shows where full precision also fits single-GPU.
VRAM breakdown →
VRAM breakdown →
VRAM breakdown →
VRAM breakdown →
VRAM breakdown →
Specifications
The A10G trades peak compute for a 150W inference envelope. 24 GB handles 13B INT4 with KV-cache headroom and 7B FP16 comfortably; 30B INT4 (≈16 GB) fits but leaves a smaller context budget than 48GB-class cards. It has Ampere tensor cores (no FP8) and no NVLink — tensor parallelism is not part of the deployment model. Bandwidth of 0.62 TB/s is roughly two-thirds of an RTX 3090, so decode throughput on the same model is correspondingly lower. No rate rows are tracked for the A10G in OpenGPU Radar's provider table.
Next steps