CodingDenseContext: 125K
Qwen 2.5 Coder 32B
Comprehensive deployment profile and benchmark telemetry for Qwen 2.5 Coder 32B. Self-hosting memory footprints, verified token economics, and direct API endpoints.
Verified Engineering Benchmarks
SWE-bench
49.2%
LiveCodeBench
57.2%
MATH-500
97.3%
MMLU
82.3%
VRAM Requirements & Sizing
๐ Quick Actions
๐ฅ๏ธ Compatible GPUs & Self-Host Pricing
Deterministic VRAM math from entity graph. Green = fits in single GPU.
| GPU | VRAM | FP16 | FP8 | INT4 | Calculator Link |
|---|---|---|---|---|---|
| H100 SXM5 | 80 GB | โ fits | โ fits | โ fits | Pre-filled โ |
| H200 | 141 GB | โ fits | โ fits | โ fits | Pre-filled โ |
| B200 | 192 GB | โ fits | โ fits | โ fits | Pre-filled โ |
| A100 80GB | 80 GB | โ fits | โ fits | โ fits | Pre-filled โ |
| L40S | 48 GB | OOM | โ fits | โ fits | Pre-filled โ |
| RTX 4090 | 24 GB | OOM | OOM | โ fits | Pre-filled โ |
โก Verified Free API Available
KV-Cache Memory Scaling
| Precision | Weights | 4K Context | 32K Context | 128K Context | Total 4K | Total 32K | Total 128K |
|---|---|---|---|---|---|---|---|
| FP16 | 65 GB | 1 GB | 4 GB | 16 GB | ~66 GB | ~69 GB | ~81 GB |
| FP8 | 33 GB | 0.5 GB | 2 GB | 8 GB | ~33.5 GB | ~35 GB | ~41 GB |
| INT4/AWQ | 18 GB | 0.25 GB | 1 GB | 4 GB | ~18.3 GB | ~19 GB | ~22 GB |
โ OptimizedWeights + KV-Cache formula:
Total = Weights + 2 ร Layers ร Heads ร Head_Dim ร Context ร Bytes_Per_Element + 2 GB CUDAProduction Run Commands
vLLM ServerCopy via DevTools โ
python3 -m vllm.entrypoints.openai.api_server \ --model Qwen/Qwen2.5-Coder-32B-Instruct \ --max-model-len 32768 \ --gpu-memory-utilization 0.95 \ --tensor-parallel-size 1 \ --quantization awq
Tensor parallelism: TP=1 ยท Max model len: 32768 ยท Quantization: AWQ
Ollama ModelfileCopy via DevTools โ
FROM Qwen/Qwen2.5-Coder-32B-Instruct
PARAMETER num_ctx 4096
PARAMETER num_thread 4
TEMPLATE "<|im_start|>{{user}}<|im_end|>
<|im_start|>assistant<|im_end|>
{{.}}"num_ctx tuned for 4096 tokens ยท Increase num_ctx for longer context (up to 128k for supported GPUs)
Host-vs-API Breakeven Analysis
Self-Host Cost
$1.19/hr
1x L40S 48GB or 1x A100 80GB ยท 320K tokens/hr
API Cost
$0.20/$M in
$0.60/$M out ยท Groq/DeepInfra/OpenRouter
Breakeven Point
3.0M tokens/day
Self-host cheaper above this volume
๐ก Breakeven Insight
At $1.19/hr self-hosting on 1x L40S 48GB or 1x A100 80GB, you break even after 3.0M input tokens per day compared to $0.20/M API costs. Below that threshold, serverless APIs (Groq , DeepInfra) are strictly cheaper. Above it, self-hosting with vLLM on spot GPUs provides better unit economics at scale.
Active API Providers & Free Tiers
| Provider | Input / 1M | Output / 1M | Speed | Free Tier Status |
|---|---|---|---|---|
| DeepInfra | $0.14 | $0.18 | 55 tok/s | Paid API Only |
| Together AI | $0.20 | $0.60 | 50 tok/s | Paid API Only |
| OpenRouter | $0.12 | $0.18 | 48 tok/s | โก 20 RPM |
| NVIDIA NIM | $0.10 | $0.15 | 500 tok/s | Paid API Only |
| GitHub Models | $0.15 | $0.20 | 30 tok/s | โก 15 RPM / 150 RPD |
OpenAI SDK Drop-in
import OpenAI from "openai";
const client = new OpenAI({
baseURL: "https://api.deepinfra.com/v1",
apiKey: process.env.DEEPINFRA_API_KEY,
});
const response = await client.chat.completions.create({
model: "qwen-2.5-coder-32b",
messages: [{ role: "user", content: "Hello" }],
});