Spec & Pricing Reference

NVIDIA A10G Cloud Pricing & Specs (2026)

The A10G is Ampere's cloud-inference SKU: 24 GB GDDR6 at 0.62 TB/s in a 150W envelope. It is the GPU behind AWS g5 instances and similar inference-optimized fleets, tuned for 7B-13B INT4 serving rather than training.

Target workload: Cloud inference for 7B-13B models & graphics rendering

Memory24GB GDDR6
Bandwidth0.62 TB/s
FP8 TFLOPSNot supported
Observed rateNo tracked rows
24 GB VRAM
Provider PublishedHIGH
SourceOpenGPU Radar telemetry row (gpu-pricing.json id: a10g)
VerifiedSep 30, 2026
Value
24 GB
Methodology

Spec fields mirrored from the verified telemetry row; board attributes per NVIDIA A10G datasheet. No rate rows tracked.

Refreshed daily from provider APIs and market scrapingSep 30, 2026
0.62 TB/s
Manufacturer SpecHIGH
SourceNVIDIA specification
VerifiedSep 30, 2026
Value
0.62 TB/s GB/s
Methodology

Peak memory bandwidth from the manufacturer specification.

Refreshed daily from provider APIs and market scrapingSep 30, 2026
Methodology →

NVIDIA A10G: Key Numbers at a Glance

Rental cost: no verified rate rows are tracked yet — this page asserts specifications only, never a price.

70B fit: Does not fit Llama 3.3 70B single-GPU — INT4 alone needs 40 GB vs 24 GB available. 7B-13B class is the single-GPU target.

Bottleneck: Memory-bandwidth bound — 0.62 TB/s is the token-generation ceiling

Specs Only — No Rate Rows Tracked
Pricing coverage for NVIDIA A10G

OpenGPU Radar does not currently track this GPU rows in data/providers.json. No price is asserted on this page — the specifications below are manufacturer-sourced.

Compare GPUs with tracked rates →

Compatible Models for NVIDIA A10G

Models from the VRAM registry whose minimum INT4 footprint (weights + KV-cache + runtime overhead) fits 24 GB. FP16 shows where full precision also fits single-GPU.

Browse all model VRAM pages →

Specifications

ArchitectureAmpere GA102 — 8nm Samsung
Memory24GB GDDR6
Bandwidth0.62 TB/s
InterconnectPCIe 4.0 (64 GB/s)
TDP150W
FP16 TFLOPS31.2
FP8 TFLOPSN/A — no FP8 Tensor Cores (Ampere generation)
FP4 TFLOPSN/A
Recommended quantizationINT8 / INT4 (GPTQ/AWQ/GGUF)
Best cluster topologyPCIe single node — inference appliance

The A10G trades peak compute for a 150W inference envelope. 24 GB handles 13B INT4 with KV-cache headroom and 7B FP16 comfortably; 30B INT4 (≈16 GB) fits but leaves a smaller context budget than 48GB-class cards. It has Ampere tensor cores (no FP8) and no NVLink — tensor parallelism is not part of the deployment model. Bandwidth of 0.62 TB/s is roughly two-thirds of an RTX 3090, so decode throughput on the same model is correspondingly lower. No rate rows are tracked for the A10G in OpenGPU Radar's provider table.

Next steps

Related GPUs

Frequently Asked Questions

How much does it cost to rent NVIDIA A10G per hour?▾
OpenGPU Radar does not currently track rate rows for NVIDIA A10G. The specifications on this page are manufacturer-sourced; no price is asserted.
Can NVIDIA A10G run 70B parameter LLMs?▾
Does not fit Llama 3.3 70B single-GPU — INT4 alone needs 40 GB vs 24 GB available. 7B-13B class is the single-GPU target.
What is NVIDIA A10G's memory bandwidth?▾
NVIDIA A10G has 0.62 TB/s of memory bandwidth across 24GB GDDR6. Token generation is memory-bandwidth bound at batch size 1, so decode throughput scales with this figure — no tokens/sec value is claimed without a benchmark row.
Specifications only — no rate rows tracked for this GPU.Spec sources: OpenGPU Radar telemetry row (gpu-pricing.json id: a10g).Methodology →