โšกUnder $0.50/hr๐Ÿง VRAM Estimatorโš–Compare GPUs๐ŸŽFree LLM APIs๐ŸŽฏModel Index
CodingDenseContext: 125K

Qwen 2.5 Coder 32B

Comprehensive deployment profile and benchmark telemetry for Qwen 2.5 Coder 32B. Self-hosting memory footprints, verified token economics, and direct API endpoints.

Verified Engineering Benchmarks

SWE-bench
49.2%
LiveCodeBench
57.2%
MATH-500
97.3%
MMLU
82.3%

VRAM Requirements & Sizing

FP16 Weights
64 GB
INT4 / GGUF (Quantized)
20 GB
Recommended GPU
1x A100
๐Ÿ”— Quick Actions
๐Ÿ–ฅ๏ธ Compatible GPUs & Self-Host Pricing

Deterministic VRAM math from entity graph. Green = fits in single GPU.

GPUVRAMFP16FP8INT4Calculator Link
H100 SXM580 GBโœ“ fitsโœ“ fitsโœ“ fitsPre-filled โ†’
H200141 GBโœ“ fitsโœ“ fitsโœ“ fitsPre-filled โ†’
B200192 GBโœ“ fitsโœ“ fitsโœ“ fitsPre-filled โ†’
A100 80GB80 GBโœ“ fitsโœ“ fitsโœ“ fitsPre-filled โ†’
L40S48 GBOOMโœ“ fitsโœ“ fitsPre-filled โ†’
RTX 409024 GBOOMOOMโœ“ fitsPre-filled โ†’
โšก Verified Free API Available
GroqNo Card
30 RPM, 14,400 requests/day via Groq Console

KV-Cache Memory Scaling

PrecisionWeights4K Context32K Context128K ContextTotal 4KTotal 32KTotal 128K
FP1665 GB1 GB4 GB16 GB~66 GB~69 GB~81 GB
FP833 GB0.5 GB2 GB8 GB~33.5 GB~35 GB~41 GB
INT4/AWQ18 GB0.25 GB1 GB4 GB~18.3 GB~19 GB~22 GB
โœ“ OptimizedWeights + KV-Cache formula: Total = Weights + 2 ร— Layers ร— Heads ร— Head_Dim ร— Context ร— Bytes_Per_Element + 2 GB CUDA

Production Run Commands

vLLM ServerCopy via DevTools โ†’
python3 -m vllm.entrypoints.openai.api_server \
  --model Qwen/Qwen2.5-Coder-32B-Instruct \
  --max-model-len 32768 \
  --gpu-memory-utilization 0.95 \
  --tensor-parallel-size 1 \
  --quantization awq

Tensor parallelism: TP=1 ยท Max model len: 32768 ยท Quantization: AWQ

Ollama ModelfileCopy via DevTools โ†’
FROM Qwen/Qwen2.5-Coder-32B-Instruct

PARAMETER num_ctx 4096
PARAMETER num_thread 4

TEMPLATE "<|im_start|>{{user}}<|im_end|>
<|im_start|>assistant<|im_end|>
{{.}}"

num_ctx tuned for 4096 tokens ยท Increase num_ctx for longer context (up to 128k for supported GPUs)

Host-vs-API Breakeven Analysis

Self-Host Cost
$1.19/hr
1x L40S 48GB or 1x A100 80GB ยท 320K tokens/hr
API Cost
$0.20/$M in
$0.60/$M out ยท Groq/DeepInfra/OpenRouter
Breakeven Point
3.0M tokens/day
Self-host cheaper above this volume
๐Ÿ’ก Breakeven Insight

At $1.19/hr self-hosting on 1x L40S 48GB or 1x A100 80GB, you break even after 3.0M input tokens per day compared to $0.20/M API costs. Below that threshold, serverless APIs (Groq , DeepInfra) are strictly cheaper. Above it, self-hosting with vLLM on spot GPUs provides better unit economics at scale.

Active API Providers & Free Tiers

ProviderInput / 1MOutput / 1MSpeedFree Tier Status
DeepInfra$0.14$0.1855 tok/sPaid API Only
Together AI$0.20$0.6050 tok/sPaid API Only
OpenRouter$0.12$0.1848 tok/sโšก 20 RPM
NVIDIA NIM$0.10$0.15500 tok/sPaid API Only
GitHub Models$0.15$0.2030 tok/sโšก 15 RPM / 150 RPD

OpenAI SDK Drop-in

import OpenAI from "openai";

const client = new OpenAI({
  baseURL: "https://api.deepinfra.com/v1",
  apiKey: process.env.DEEPINFRA_API_KEY,
});

const response = await client.chat.completions.create({
  model: "qwen-2.5-coder-32b",
  messages: [{ role: "user", content: "Hello" }],
});

Related Engineering Guides & Analysis

What should I do next?