Research Workflows: Which LLM & GPU Fits Your Document Analysis?
Sizing local LLMs for research, RAG, and long-document analysis. Correct VRAM math for Llama 3.3 70B, DeepSeek R1, and Phi-4 at long context, plus a task-to-GPU fit matrix.
Direct answer
For LLM-based research and document analysis (RAG, long-context summarization, code-as-tools), the GPU you need depends on your context length, KV-cache precision, and whether you quantize:
- Llama 3.3 70B at 128K context: โ68 GB total at INT4 weights + FP8 KV (or โ90 GB with FP16 KV) โ fits 1ร H200 (141 GB); does not fit a single H100 (80 GB) or any consumer card at full residency.
- DeepSeek R1 (671B) at 128K context: โ376 GB at INT4 (weights 336 GB + MLA KV ~4 GB) โ 3ร H200 minimum, 4ร recommended for even MoE expert split; โ745 GB at FP8 โ 6ร H200 or 4ร B200. Never a single GPU.
- Phi-4 14B at 16K context: โ10 GB at INT4 โ fits a single RTX 4090 (24 GB) โ viable for local research notes and agent loops.
- 70B-class models on a single consumer 24 GB card: only feasible with CPU offload or aggressive 2-bit quants at short context โ see Can I run a 70B LLM on a 24GB GPU? for the feasibility breakdown.
For document analysis under ~100 pages per run, start at INT4 on a consumer GPU or a single rental, and move to datacenter cards only when you regularly need 70B-class quality at 32Kโ128K context.
Context is the hidden cost
Research workflows are context-bound. A 300-page PDF at a typical ~350โ500 tokens per page is ~105Kโ150K tokens before any conversation history, system prompt, or retrieved snippets are added โ which is why long-document work quickly reaches 128K-context territory. At 128K context, the KV-cache on a 70B GQA model is ~20 GB at FP8 KV or ~40 GB at FP16 KV (see KV-cache explainer) โ often as large as the quantized weights themselves.
The total VRAM equation is:
total = (model_weights + kv_cache + cuda_overhead + activations) ร 1.10
Where the 1.10 multiplier covers ~10% fragmentation headroom. The cuda_overhead constant is 1.2 GB; activations add 0.4 GB (batch=1). The standard GQA KV-cache term:
kv_cache (GB) = 2 ร num_layers ร num_kv_heads ร head_dim ร context_length ร bytes_per_element รท 1,073,741,824
For Llama 3.3 70B use 80 layers (80 transformer layers, 8 KV heads, head_dim 128 โ the same architecture as Llama 3.1 70B; do not use the repository registry's incorrect 28-layer figure). At 128K context:
- FP8 KV (1 byte):
2 ร 80 ร 8 ร 128 ร 128,000 ร 1 รท 2ยณโฐ โ 20 GB - FP16 KV (2 bytes):
โ 40 GB
DeepSeek R1 is different: Multi-Latent Attention stores only kv_lora_rank (512) + qk_rope_head_dim (64) = 576 dims per token per layer โ at 61 layers and 128K context that is only ~4.2 GB, regardless of weight precision. R1's binding constraint is always weights (336 GB INT4 / 671 GB FP8), never KV-cache.
Model & VRAM sizing for research tasks
All totals below: weights + KV at 128K + 1.6 GB overhead, ร 1.10 headroom, batch=1. Weight figures from the model registry (minVramInt4Gb / minVramFp16Gb); KV from the formulas above with correct layer counts (Llama 3.3 70B = 80 layers; DeepSeek R1 = 61 layers MLA).
| Model | Precision | Weights | KV @128K | Total @128K | GPU fit |
|---|---|---|---|---|---|
| Llama 3.3 70B | INT4 + FP8 KV | 40 GB | ~20 GB | ~68 GB | 1ร H200 SXM (141 GB) |
| Llama 3.3 70B | INT4 + FP16 KV | 40 GB | ~40 GB | ~90 GB | 1ร H200 SXM |
| Llama 3.3 70B | FP8 + FP8 KV | 71 GB | ~20 GB | ~102 GB | 1ร H200 SXM |
| Llama 3.3 70B | FP8 + FP16 KV | 71 GB | ~40 GB | ~124 GB | 1ร H200 SXM (tight) |
| DeepSeek R1 671B | INT4 (MLA KV) | 336 GB | ~4.2 GB | ~376 GB | 3ร H200 min / 4ร H200 rec. |
| DeepSeek R1 671B | FP8 (MLA KV) | 671 GB | ~4.2 GB | ~745 GB | 6ร H200 or 4ร B200 |
| Phi-4 14B | INT4 | 10 GB | ~1 GB @16K | ~12 GB @16K | RTX 4090 (24 GB) |
| Qwen 2.5 72B | INT4 + FP8 KV | 40 GB | ~20 GB | ~68 GB | 1ร H200 SXM |
Why not 4ร H100 for R1 INT4? 4ร H100 SXM = 320 GB < 376 GB needed โ it does not fit. INT4 R1 needs either 5ร H100 (400 GB) or 4ร H200 (564 GB; recommended for the 64รท4 expert split). FP8 R1 cannot fit any single GPU: even a B200 (192 GB) is far below the 671 GB of FP8 weights. See the DeepSeek R1 serving guide for cluster configs.
Pick precision by GPU: consumer cards (RTX 4090, L40S) need INT4 to fit anything above ~30B-class models โ and even then only at short context. Datacenter cards (H100, H200, B200) support FP8 losslessly and only need INT4 for the largest models (DeepSeek R1, Llama 3.1 405B). Quantization formats explained.
Research workload fit matrix
| Task | Model class | Context needed | Precision | Recommended |
|---|---|---|---|---|
| <5 docs, short summarization | 8โ14B (Phi-4, Gemma) | 4Kโ16K | INT4 | RTX 4090 (24 GB) |
| <20 docs, chunked retrieval (RAG) | 14โ32B (Qwen 2.5 14B/32B) | 8Kโ32K | INT4 | 1ร RTX 4090 or L40S |
| 20โ100 docs, long-context reasoning | 70B dense (Llama 3.3) | 32Kโ128K | INT4 | 1ร H200 (or 2ร RTX 4090 offloaded โ see 24GB feasibility) |
| Agent loops + tool calls | 30โ70B | 32Kโ128K effective | INT4 | 1ร H200 or 2ร RTX 4090; plan KV headroom for growing conversation state |
| DeepSeek R1 (671B) full-context | 671B MoE | 128K | INT4 | 4ร H200 SXM5 (TP=4) |
Agent loops: state is sequential, not duplicated โ you do not need 2ร the weights. What grows across turns is the KV-cache (each tool result and reply appends tokens). Budget headroom for a longer effective context (e.g. 32K working set on a model sized for 8K), not double the model.
Practical GPU selection by budget
- Budget (local, ~$1,600 capex or $0.34โ0.74/hr rental): single RTX 4090/5080 (16โ24 GB) with INT4 โ fully covers โค32B models and Phi-4-class work; 70B requires offload/2-bit (see 24GB guide). Buying advice: best consumer GPU for AI 2026.
- Single datacenter card (~$1.89โ2.79/hr): H100 SXM5 (80 GB) fits 70B INT4 at โค32K comfortably; H200 SXM5 (141 GB) is the single-card answer for 70B at full 128K (see H200 128K FP8 config). Compare H100 vs H200 ROI.
- Multi-GPU cluster (>4 GPUs): required for DeepSeek R1 671B or FP16 serving of 70B at 128K. NVLink bandwidth becomes the binding constraint at 3+ GPUs โ B200 NVLink-5 scales better than H100 NVLink-4. B200 NVLink scaling analysis.
Spot vs on-demand: research workloads with retry-on-failure semantics are ideal for spot instances. The cloud GPU pricing index tracks hourly rates across providers; the breakeven math for API vs self-hosting shows when renting beats calling an API.
Worked example: sizing Llama 3.3 70B for a 128K research job
- Model weights: 70.6B parameters ร 0.5 bytes (INT4 floor) = 35.3 GB; registry rounds to 40 GB including projection buffers (real GGUF Q4_K_M files are ~42.5 GB โ size from actual files for local serving).
- KV-cache (128K, 80 layers, 8 KV heads, FP8 KV):
2 ร 80 ร 8 ร 128 ร 128,000 ร 1 รท 2ยณโฐ โ 20 GB. - Overhead: 1.2 GB CUDA + 0.4 GB activations.
- Subtotal: 40 + 20 + 1.6 = 61.6 GB. Apply 10% headroom โ ~68 GB โ fits 1ร H200 (141 GB) with room for batching.
With the runtime defaulting KV-cache to FP16 (common in llama.cpp/Ollama unless you opt into quantized KV), KV doubles to ~40 GB and the total becomes ~90 GB โ still H200-class, no longer H100-class. Always confirm your runtime's KV precision setting.
Limitations & common mistakes
- Forgetting KV precision: many local runtimes (llama.cpp, Ollama) default KV-cache to FP16 even when weights are INT4/Q4_K_M. At long context the KV-cache, not the weights, is often the binding constraint.
- Wrong layer counts: Llama 3.3 70B has 80 layers (same as 3.1 70B). Repository registry entries with 28 or 32 layers will under-provision KV-cache by 2.5โ3ร. DeepSeek R1 has 61 layers with MLA โ use 576 dims/token, not GQA head math.
- Assuming linear scaling: batch size and multi-GPU tensor parallelism incur overhead. Two H100s do not double effective capacity after NVLink and TP overhead (~10โ15%).
- Context vs. output tokens: 128K context window โ 128K output. Long generated summaries grow KV-cache during decode โ monitor utilization, not just load-time fit.
- MLA compression is not universal: only DeepSeek R1/V3 and a few 2026 models use compressed KV-cache. Standard GQA models (Llama, Qwen, Phi) use the full KV formula.
FAQ
Can I run a 70B research model on a single GPU? On datacenter single cards: yes โ 1ร H200 at INT4 or FP8 covers 70B at 128K (see above). On a single H100 (80 GB): INT4 works at โค32K context, not full 128K. On a single 24 GB consumer card: only via CPU offload or 2-bit quants at short context โ read the 24GB feasibility guide before buying.
Is INT4 worth the quality loss for research? For retrieval-augmented workflows where the LLM summarizes or reasons over retrieved text, community evaluations typically find 4-bit K-quants within a few points of FP8/BF16 on QA-style benchmarks โ but reasoning-heavy math/coding tasks degrade more. Run your own evaluation set; we do not publish a single universal delta.
What about the KV-cache trap at long context? At 128K context, KV-cache can consume 20โ40+ GB on a 70B GQA model โ comparable to the INT4 weights. See KV-cache explained for the formula, flash-attention mitigations, and the MLA exception for DeepSeek models.
Sources
- Layer counts and MLA dims: HuggingFace
config.jsonfor DeepSeek-R1 (num_hidden_layers=61,kv_lora_rank=512,qk_rope_head_dim=64); Llama 3.3 70B = 80 layers (Meta model card / transformersconfig.json), 8 KV heads, head_dim 128 - Weight VRAM:
src/data/models-registry.json(minVramInt4Gb,minVramFp16Gb); GGUF file sizes from 24GB feasibility guide - KV-cache formula: KV-cache explainer; totals cross-checked against the deterministic engine in LLM VRAM calculator guide
- Cluster sizing: DeepSeek R1 serving guide, FP8 vs INT4 memory analysis
- Pricing: cheapest cloud GPUs โ verify live rates before procurement
- Data verification date: 2026-10-11