Hardware Sizing2026-10-10โ€ขBy Sreeโ€ข5 min read

Research Workflows: Which LLM & GPU Fits Your Document Analysis?

Sizing local LLMs for research, RAG, and long-document analysis. Correct VRAM math for Llama 3.3 70B, DeepSeek R1, and Phi-4 at long context, plus a task-to-GPU fit matrix.

Direct answer

For LLM-based research and document analysis (RAG, long-context summarization, code-as-tools), the GPU you need depends on your context length, KV-cache precision, and whether you quantize:

  • Llama 3.3 70B at 128K context: โ‰ˆ68 GB total at INT4 weights + FP8 KV (or โ‰ˆ90 GB with FP16 KV) โ€” fits 1ร— H200 (141 GB); does not fit a single H100 (80 GB) or any consumer card at full residency.
  • DeepSeek R1 (671B) at 128K context: โ‰ˆ376 GB at INT4 (weights 336 GB + MLA KV ~4 GB) โ€” 3ร— H200 minimum, 4ร— recommended for even MoE expert split; โ‰ˆ745 GB at FP8 โ€” 6ร— H200 or 4ร— B200. Never a single GPU.
  • Phi-4 14B at 16K context: โ‰ˆ10 GB at INT4 โ€” fits a single RTX 4090 (24 GB) โ€” viable for local research notes and agent loops.
  • 70B-class models on a single consumer 24 GB card: only feasible with CPU offload or aggressive 2-bit quants at short context โ€” see Can I run a 70B LLM on a 24GB GPU? for the feasibility breakdown.

For document analysis under ~100 pages per run, start at INT4 on a consumer GPU or a single rental, and move to datacenter cards only when you regularly need 70B-class quality at 32Kโ€“128K context.

Context is the hidden cost

Research workflows are context-bound. A 300-page PDF at a typical ~350โ€“500 tokens per page is ~105Kโ€“150K tokens before any conversation history, system prompt, or retrieved snippets are added โ€” which is why long-document work quickly reaches 128K-context territory. At 128K context, the KV-cache on a 70B GQA model is ~20 GB at FP8 KV or ~40 GB at FP16 KV (see KV-cache explainer) โ€” often as large as the quantized weights themselves.

The total VRAM equation is:

total = (model_weights + kv_cache + cuda_overhead + activations) ร— 1.10

Where the 1.10 multiplier covers ~10% fragmentation headroom. The cuda_overhead constant is 1.2 GB; activations add 0.4 GB (batch=1). The standard GQA KV-cache term:

kv_cache (GB) = 2 ร— num_layers ร— num_kv_heads ร— head_dim ร— context_length ร— bytes_per_element รท 1,073,741,824

For Llama 3.3 70B use 80 layers (80 transformer layers, 8 KV heads, head_dim 128 โ€” the same architecture as Llama 3.1 70B; do not use the repository registry's incorrect 28-layer figure). At 128K context:

  • FP8 KV (1 byte): 2 ร— 80 ร— 8 ร— 128 ร— 128,000 ร— 1 รท 2ยณโฐ โ‰ˆ 20 GB
  • FP16 KV (2 bytes): โ‰ˆ 40 GB

DeepSeek R1 is different: Multi-Latent Attention stores only kv_lora_rank (512) + qk_rope_head_dim (64) = 576 dims per token per layer โ€” at 61 layers and 128K context that is only ~4.2 GB, regardless of weight precision. R1's binding constraint is always weights (336 GB INT4 / 671 GB FP8), never KV-cache.

Model & VRAM sizing for research tasks

All totals below: weights + KV at 128K + 1.6 GB overhead, ร— 1.10 headroom, batch=1. Weight figures from the model registry (minVramInt4Gb / minVramFp16Gb); KV from the formulas above with correct layer counts (Llama 3.3 70B = 80 layers; DeepSeek R1 = 61 layers MLA).

ModelPrecisionWeightsKV @128KTotal @128KGPU fit
Llama 3.3 70BINT4 + FP8 KV40 GB~20 GB~68 GB1ร— H200 SXM (141 GB)
Llama 3.3 70BINT4 + FP16 KV40 GB~40 GB~90 GB1ร— H200 SXM
Llama 3.3 70BFP8 + FP8 KV71 GB~20 GB~102 GB1ร— H200 SXM
Llama 3.3 70BFP8 + FP16 KV71 GB~40 GB~124 GB1ร— H200 SXM (tight)
DeepSeek R1 671BINT4 (MLA KV)336 GB~4.2 GB~376 GB3ร— H200 min / 4ร— H200 rec.
DeepSeek R1 671BFP8 (MLA KV)671 GB~4.2 GB~745 GB6ร— H200 or 4ร— B200
Phi-4 14BINT410 GB~1 GB @16K~12 GB @16KRTX 4090 (24 GB)
Qwen 2.5 72BINT4 + FP8 KV40 GB~20 GB~68 GB1ร— H200 SXM

Why not 4ร— H100 for R1 INT4? 4ร— H100 SXM = 320 GB < 376 GB needed โ€” it does not fit. INT4 R1 needs either 5ร— H100 (400 GB) or 4ร— H200 (564 GB; recommended for the 64รท4 expert split). FP8 R1 cannot fit any single GPU: even a B200 (192 GB) is far below the 671 GB of FP8 weights. See the DeepSeek R1 serving guide for cluster configs.

Pick precision by GPU: consumer cards (RTX 4090, L40S) need INT4 to fit anything above ~30B-class models โ€” and even then only at short context. Datacenter cards (H100, H200, B200) support FP8 losslessly and only need INT4 for the largest models (DeepSeek R1, Llama 3.1 405B). Quantization formats explained.

Research workload fit matrix

TaskModel classContext neededPrecisionRecommended
<5 docs, short summarization8โ€“14B (Phi-4, Gemma)4Kโ€“16KINT4RTX 4090 (24 GB)
<20 docs, chunked retrieval (RAG)14โ€“32B (Qwen 2.5 14B/32B)8Kโ€“32KINT41ร— RTX 4090 or L40S
20โ€“100 docs, long-context reasoning70B dense (Llama 3.3)32Kโ€“128KINT41ร— H200 (or 2ร— RTX 4090 offloaded โ€” see 24GB feasibility)
Agent loops + tool calls30โ€“70B32Kโ€“128K effectiveINT41ร— H200 or 2ร— RTX 4090; plan KV headroom for growing conversation state
DeepSeek R1 (671B) full-context671B MoE128KINT44ร— H200 SXM5 (TP=4)

Agent loops: state is sequential, not duplicated โ€” you do not need 2ร— the weights. What grows across turns is the KV-cache (each tool result and reply appends tokens). Budget headroom for a longer effective context (e.g. 32K working set on a model sized for 8K), not double the model.

Practical GPU selection by budget

  • Budget (local, ~$1,600 capex or $0.34โ€“0.74/hr rental): single RTX 4090/5080 (16โ€“24 GB) with INT4 โ€” fully covers โ‰ค32B models and Phi-4-class work; 70B requires offload/2-bit (see 24GB guide). Buying advice: best consumer GPU for AI 2026.
  • Single datacenter card (~$1.89โ€“2.79/hr): H100 SXM5 (80 GB) fits 70B INT4 at โ‰ค32K comfortably; H200 SXM5 (141 GB) is the single-card answer for 70B at full 128K (see H200 128K FP8 config). Compare H100 vs H200 ROI.
  • Multi-GPU cluster (>4 GPUs): required for DeepSeek R1 671B or FP16 serving of 70B at 128K. NVLink bandwidth becomes the binding constraint at 3+ GPUs โ€” B200 NVLink-5 scales better than H100 NVLink-4. B200 NVLink scaling analysis.

Spot vs on-demand: research workloads with retry-on-failure semantics are ideal for spot instances. The cloud GPU pricing index tracks hourly rates across providers; the breakeven math for API vs self-hosting shows when renting beats calling an API.

Worked example: sizing Llama 3.3 70B for a 128K research job

  1. Model weights: 70.6B parameters ร— 0.5 bytes (INT4 floor) = 35.3 GB; registry rounds to 40 GB including projection buffers (real GGUF Q4_K_M files are ~42.5 GB โ€” size from actual files for local serving).
  2. KV-cache (128K, 80 layers, 8 KV heads, FP8 KV): 2 ร— 80 ร— 8 ร— 128 ร— 128,000 ร— 1 รท 2ยณโฐ โ‰ˆ 20 GB.
  3. Overhead: 1.2 GB CUDA + 0.4 GB activations.
  4. Subtotal: 40 + 20 + 1.6 = 61.6 GB. Apply 10% headroom โ†’ ~68 GB โ†’ fits 1ร— H200 (141 GB) with room for batching.

With the runtime defaulting KV-cache to FP16 (common in llama.cpp/Ollama unless you opt into quantized KV), KV doubles to ~40 GB and the total becomes ~90 GB โ€” still H200-class, no longer H100-class. Always confirm your runtime's KV precision setting.

Limitations & common mistakes

  • Forgetting KV precision: many local runtimes (llama.cpp, Ollama) default KV-cache to FP16 even when weights are INT4/Q4_K_M. At long context the KV-cache, not the weights, is often the binding constraint.
  • Wrong layer counts: Llama 3.3 70B has 80 layers (same as 3.1 70B). Repository registry entries with 28 or 32 layers will under-provision KV-cache by 2.5โ€“3ร—. DeepSeek R1 has 61 layers with MLA โ€” use 576 dims/token, not GQA head math.
  • Assuming linear scaling: batch size and multi-GPU tensor parallelism incur overhead. Two H100s do not double effective capacity after NVLink and TP overhead (~10โ€“15%).
  • Context vs. output tokens: 128K context window โ‰  128K output. Long generated summaries grow KV-cache during decode โ€” monitor utilization, not just load-time fit.
  • MLA compression is not universal: only DeepSeek R1/V3 and a few 2026 models use compressed KV-cache. Standard GQA models (Llama, Qwen, Phi) use the full KV formula.

FAQ

Can I run a 70B research model on a single GPU? On datacenter single cards: yes โ€” 1ร— H200 at INT4 or FP8 covers 70B at 128K (see above). On a single H100 (80 GB): INT4 works at โ‰ค32K context, not full 128K. On a single 24 GB consumer card: only via CPU offload or 2-bit quants at short context โ€” read the 24GB feasibility guide before buying.

Is INT4 worth the quality loss for research? For retrieval-augmented workflows where the LLM summarizes or reasons over retrieved text, community evaluations typically find 4-bit K-quants within a few points of FP8/BF16 on QA-style benchmarks โ€” but reasoning-heavy math/coding tasks degrade more. Run your own evaluation set; we do not publish a single universal delta.

What about the KV-cache trap at long context? At 128K context, KV-cache can consume 20โ€“40+ GB on a 70B GQA model โ€” comparable to the INT4 weights. See KV-cache explained for the formula, flash-attention mitigations, and the MLA exception for DeepSeek models.

Sources