AI Video Generation 24/7: GPU Requirements, Throughput and Monthly Cost
What hardware and budget run an automated AI video pipeline continuously — Wan 2.2, LTX-2, HunyuanVideo VRAM needs, clips-per-day math, cloud vs local electricity costs, and queue operations.
Direct answer
A continuous AI video-generation pipeline is a queue-and-worker system, not a screensaver. One warm GPU worker on a 24 GB card (RTX 4090-class) can produce roughly 120–250 finished 5-second 720p clips per day depending on the model — LTX-Video at the fast end (~720/day), Wan 2.2 14B at the quality end (~249/day), at a labeled 75% utilization assumption. Renting that GPU runs about $250/month at a dated $0.34/hr snapshot; owning it costs roughly $64/month in electricity plus amortized hardware, with a ~9-month buy-back against rental at sustained utilization.
The short version:
| Question | Answer | Basis |
|---|---|---|
| Minimum serious VRAM | 12 GB (LTX-Video tier); 24 GB for quality models (Wan 2.2 14B FP8, HunyuanVideo 1.5) | Community-documented model requirements, mid-2026 |
| Best starter model | Wan 2.2 TI2V-5B or LTX-Video 0.9.5 — both fit 12–24 GB | Apache-2.0 / free-under-$10M licenses |
| Best quality model that fits 24 GB | Wan 2.2 A14B (27B MoE, 14B active) at FP8/GGUF with T5 offload | ~6–8 GB quantized; 54–65 GB at FP16 |
| Starter config | 1× RTX 4090 (24 GB), ~$250/mo rented or ~$64/mo electricity owned | Cloud rental page |
| Higher-throughput config | 4× RTX 4090 workers, or 1× L40S (48 GB) for fewer OOM retries | Linear-ish scaling if queue is deep |
| Local cheaper than cloud when | Utilization sustains >~45% and you keep the box 3+ years | Same framework as cloud vs self-hosting |
No configuration here promises uninterrupted production. GPUs thermally throttle, OOM, and crash; models ship breaking updates; queues stall. What follows is the arithmetic with the assumptions exposed so you can re-run it.
What a 24/7 video pipeline actually is
A durable pipeline has five stages:
- Prompt intake — API, cron job, or upstream app submits jobs (text prompt, optional start image, seed, resolution, duration).
- Queue — durable broker (Redis, SQS, Postgres-backed) holding jobs with retry state. Never generate directly from the HTTP handler.
- Warm GPU worker — a long-lived process (typically ComfyUI with its HTTP API, or a custom diffusers loop) that keeps the model loaded. Cold-loading Wan 2.2 or HunyuanVideo weights takes minutes; per-job loading destroys throughput. Load once, generate thousands.
- Encode and validate — VAE decode to frames, H.264/H.265 encode, sanity checks (file size > 0, frame count matches, no NaN black frames).
- Persist and index — write clips to local NVMe, then tier to object storage; record job metadata (prompt, seed, model, duration, status) in a database for retrieval and billing.
The GPU only touches stage 3. Everything else is ordinary backend engineering — and everything else is where unattended pipelines usually fail.
Model families and workflows
All figures below are for open-weights models you can run yourself. Closed APIs (Sora, Veo, Runway) sidestep the hardware question entirely and price per clip or per second; they are the default choice under ~12 GB of local VRAM — see our local AI setup guide for the general "can my PC do this" framing.
| Family | Sizes | Workflow | License (verify before commercial use) | Standout |
|---|---|---|---|---|
| Wan 2.2 / 2.1 (Alibaba) | 1.3B, 5B TI2V, 14B MoE (27B total, 14B active) | T2V, I2V | Apache 2.0 (2.1 and 2.2) | Highest publicly scored open quality (VBench ~84.7–86.2%); runs on 8 GB+ with quantization |
| LTX-Video / LTX-2 (Lightricks) | 2B (0.9.5), 22B dual-stream (LTX-2.x) | T2V, I2V; LTX-2 adds synchronized audio | Free under $10M revenue; paid tier above | Fastest per clip — LTX-2 generates video+audio in one pass |
| HunyuanVideo 1.5 (Tencent) | 8.3B (base ~13B lineage) | T2V, I2V | Tencent Community license — excludes EU/UK/South Korea; verify territory | Cinematic look; distilled builds ~3–5 min/clip on a 4090 |
| CogVideoX (Zhipu) | 2B, 5B | T2V, I2V | Apache 2.0 (2B) | Easiest to run (~7–10 GB quantized); quality now behind the 2026 leaders |
| Mochi 1 (Genmo) | 10B | T2V 480p | Apache 2.0 | Research-grade motion; heavy (60 GB+ FP16; ~24 GB with optimizations) |
| Stable Video Diffusion (Stability) | 1.5B | I2V only | Stability community license | Lightest I2V option (~8 GB FP16); aging quality |
Workflow types:
- T2V — prompt to clip. Default for automated stock/b-roll pipelines.
- I2V — start image plus motion prompt. Dominant for product and character consistency; you need an image pipeline (e.g. FLUX — see our FLUX serving economics) upstream.
- V2V / edit — restyle or extend an existing clip. ComfyUI workflows for Wan and HunyuanVideo support this; heavier per job.
Video models are CUDA-first. NVIDIA GPUs are the practical path; ROCm/Vulkan ports exist in community projects but are slower and break more often — do not plan a 24/7 AMD pipeline on hope.
VRAM requirements
Community-documented footprints for representative models, mid-2026. Video models are diffusion transformers with large text encoders — the T5-XXL encoder alone is ~9.4B parameters and often dominates the VRAM budget. Offloading it to system RAM is the standard trick for fitting 14B-class models on 24 GB cards.
| Model | FP16 / BF16 | Quantized (FP8 / GGUF) | Practical card |
|---|---|---|---|
| Wan 2.2 A14B | 54–65 GB | 6–8 GB with T5 offload; 480p–720p workable on 24 GB | RTX 4090 (24 GB) at FP8; L40S/A100 for FP16 |
| Wan 2.2 TI2V-5B | 16–20 GB | 8–12 GB | 12–24 GB |
| Wan 2.1 T2V-1.3B | — | 8.19 GB | 8 GB+ (entry) |
| HunyuanVideo 1.5 | 24–28 GB | 14–16 GB (FP8) | 16–24 GB (license permitting) |
| HunyuanVideo (base 13B) | ~47–58 GB | ~8–22 GB depending on quant/resolution | 24 GB quantized; 48 GB+ comfortable |
| LTX-2.x | ~20 GB | 6–8 GB with tiling; official ~32 GB for full-quality 720p | 24 GB quantized; 48 GB for headroom |
| LTX-Video 0.9.5 (2B) | ~10 GB | ~4–6 GB | 12 GB — the practical floor |
| CogVideoX-5B | ~18 GB | 7–10 GB | 12 GB |
| Mochi 1 | 80 GB+ | ~16–24 GB with attention optimizations | 24 GB marginal; 48 GB+ sane |
| SVD 1.5B | ~8 GB | — | 8 GB |
Resolution, frames, duration, precision — the four levers:
- Resolution scales activation memory roughly with pixels. 480p → 720p is ~2.2× the pixels; 720p → 1080p is another ~2.2×. If you OOM, drop resolution first.
- Frame count / duration scales temporal activations. Wan's standard 5 s at 24 fps = 121 frames; HunyuanVideo uses 129 frames at 720p. Doubling duration roughly doubles temporal memory (attention implementations vary).
- Precision is the biggest single lever: FP16 → FP8 halves weights; GGUF Q4 cuts them ~4×. Quality regressions are task-specific — measure on your prompts, not on aggregate scores.
- Batch size — stay at batch 1 for unattended single-stream workers. Batching helps only if your queue is deep and VRAM headroom is large.
Our VRAM calculator models LLM footprints; video models add diffusion-specific activation and VAE costs that the LLM formula does not capture. Treat its output as a lower bound for video, then add 20–40% for the pipeline's peak (text encoder + DiT + VAE decode co-resident at some point in the graph).
Throughput: clips per day
Formula:
clips_per_day = (86400 s / seconds_per_clip) × utilization
clips_per_month = clips_per_day × 30.4
seconds_per_clip below are community measurements on an RTX 4090 at 5 s, 720p-class output — not our benchmarks, not controlled lab runs. utilization = 0.75 is our labeled assumption: real pipelines lose time to failed jobs, queue gaps, model reloads after crashes, and maintenance. If your harness measures higher sustained utilization, multiply accordingly; if lower, expect less.
| Model (RTX 4090) | Reported per-clip time | Raw clips/day (100%) | Effective/day @ 75% | Effective/month @ 75% | Source |
|---|---|---|---|---|---|
| LTX-Video 0.9.5, 5 s | ~90 s | 960 | 720 | ~21,900 | Three-reviewer head-to-head, Jul 2026 |
| Wan 2.2 14B FP8, 5 s | ~260 s (4 min 20 s) | 332 | 249 | ~7,580 | Same head-to-head |
| Wan 2.2 TI2V-5B, 5 s | ~540 s (<9 min) | 160 | 120 | ~3,650 | LocalAI Master, 2026 |
| HunyuanVideo 1.5 distilled, 5 s | ~180–300 s | 288–480 | 216–360 | ~6,600–11,000 | LocalAI Master, May 2026 |
| LTX-2.x, ~10 s | ~150 s (2.5 min) | 576 | 432 | ~13,100 | Local measurements via MagicHour, Jul 2026 |
| RTX 5090, Wan 2.1 480p | ~144 s (~25 clips/hr) | 600 | 450 | ~13,700 | Salad benchmark, Aug 2025 |
| RTX 3060 12 GB, Wan 14B, 840×420 | ~900 s (15 min / 81 frames) | 96 | 72 | ~2,200 | lilting.ch measurement, Feb 2026 |
Reading the table honestly:
- LTX-Video is 3× faster than Wan 2.2 14B per clip on the same card. If your product tolerates LTX's prompt-adherence quirks (users report needing LLM prompt rewriting), it is the throughput champion.
- The 3060 row exists to price the floor. 72 clips/day at 480p-ish output is a hobby pipeline, not a product.
- RTX 5090 roughly doubles the 4090 in the Salad 480p measurement — consistent with ~1.5–2× generational uplift, not a revolution. Do not extrapolate 5090 numbers to 720p 14B without measuring.
- A 5-second clip is 5 seconds of video. 7,580 clips/month × 5 s ≈ 10.5 hours of finished video per month from one 4090 running Wan 2.2 14B at 75% utilization. Budget accordingly — most 24/7 "video bots" are producing short social-format clips, not features.
Monthly cost model
Formulas (all inputs exposed):
cloud_monthly = hourly_rate × 730 h # 24/7 residency
electricity = (gen_hours × load_kw + idle_hours × idle_kw) × price_kwh
where gen_hours = util × 730, idle_hours = (1-util) × 730
amortization = gpu_cost / ownership_months # GPU only; host excluded
local_monthly = electricity + amortization
cost_per_clip = monthly_cost / clips_per_month
storage_monthly = (clips_per_month × avg_mb / 1e6) × $/TB # object storage
Cloud rental (dated snapshots, Sep–Oct 2026)
| GPU | Rate | Monthly (730 h) | Notes |
|---|---|---|---|
| RTX 4090 | $0.34/hr (RunPod Community) | $248 | Secure tier $0.74/hr → $540/mo; Vast.ai median ~$0.40 |
| RTX 5090 | $0.60/hr (CloudRift on-demand) | $438 | 32 GB — LTX-2 full-quality headroom |
| L40S 48 GB | $0.79/hr (RunPod Community) | $577 | More VRAM, not more speed — 864 GB/s < 4090's ~1,008 GB/s; buy it for fewer OOM retries and less quantization |
| A100 80 GB | $1.05/hr (CloudRift) | $767 | Full-precision 14B-class without quantization gymnastics |
Per-clip cloud cost at 75% utilization, RTX 4090 at $0.34/hr:
- Wan 2.2 14B: $248 ÷ 7,580 ≈ $0.033/clip
- LTX-Video: $248 ÷ 21,900 ≈ $0.011/clip
These sit in the same band as hosted per-clip APIs (Wan 2.1 720p 5 s ≈ $0.40/clip on one provider's public rate card; ~$0.22–0.72/clip ranges reported for LTX/Wan endpoints mid-2026) — self-hosting wins on unit cost only if utilization is real. At 25% utilization the same 4090 cloud box produces a quarter of the clips at the same $248, tripling per-clip cost.
Local ownership (RTX 4090 example)
Inputs (all labeled — swap your own):
- GPU purchase: $1,599 launch MSRP (Oct 2022); used-market prices vary — see the consumer GPU guide
- Ownership horizon: 36 months (input)
- System draw under generation: ~600 W (450 W GPU + ~150 W host) — measure yours
- Idle draw: ~100 W
- Utilization: 75% (same assumption as throughput)
- Electricity: $0.1834/kWh — US residential average, ~Sept 2026 (EIA-derived state tables)
gen_hours = 0.75 × 730 = 547.5 h
idle_hours = 182.5 h
kWh/month = 547.5 × 0.6 + 182.5 × 0.1 = 346.75 kWh
electricity = 346.75 × 0.1834 ≈ $63.60
amortization = 1599 / 36 ≈ $44.42
local_monthly ≈ $108
Per clip (Wan 2.2 14B): $108 ÷ 7,580 ≈ $0.014/clip — about 2.3× cheaper than the $0.34/hr cloud box, before counting the host build (add ~$500–800 one-time if you don't have one; it changes break-even by a month or two).
Break-even vs renting the same GPU continuously:
months_to_payback = gpu_cost / (cloud_monthly - local_monthly)
= 1599 / (248 - 64) ≈ 8.7 months
That number collapses if utilization drops: at 30% utilization, cloud cost for the same (smaller) output falls proportionally with usage only if you can scale the rental down — a 24/7 "always-on" cloud box does not. Owning wins when you will actually keep the queue fed for 9+ months; renting wins when demand is spiky or uncertain. The general framework (utilization, depreciation, ops overhead) is worked in cloud GPU vs self-hosting.
Storage
Measured 5 s 720p outputs: ~2.8 MB (Wan at 16 fps) to ~8.9 MB (LTX at 24 fps). Budget 5 MB average:
| Pipeline | Clips/month | Raw GB/month | Object storage @ $6.95/TB (B2) | @ $15/TB (R2) |
|---|---|---|---|---|
| Wan 2.2 14B | 7,580 | ~38 GB | ~$0.26 | ~$0.57 |
| LTX-Video | 21,900 | ~110 GB | ~$0.76 | ~$1.64 |
Object storage is noise next to the GPU bill. What is not noise: local scratch. VAE decode and temporary frame buffers want 2–3× the final file size free on fast NVMe; model checkpoints (Wan 2.2 14B GGUF ~12–15 GB + quantized T5-XXL ~5–9 GB, plus variants) want 50–100 GB. Budget a 1–2 TB NVMe for the worker box. Backblaze B2's current published rate is $6.95/TB-month with free egress up to 3× stored data; Cloudflare R2 is $0.015/GB-month ($15/TB) with free egress — both make sense for the archive tier, neither for hot scratch.
Running it unattended: queues, retries, monitoring
The GPU is the easy part. These are the failure modes that stop a 24/7 pipeline at 3 a.m.:
Job queue design
- Durable broker with at-least-once delivery and per-job retry counters. Cap retries (3 is typical); route exhausted jobs to a dead-letter queue with the full prompt + error attached.
- Idempotency keys (prompt hash + seed + model version) so a retried job doesn't silently double-bill or double-write.
- Priority lanes if interactive iteration shares the box with the overnight batch.
Retries and failure taxonomy
- CUDA OOM — dominant failure for video. Catch it, reduce resolution or frame count by one notch, retry once, then dead-letter. Do not blindly retry the same graph.
- NaN / black-frame outputs — VAE or precision bugs. Validate frame stats before persisting; requeue with a different seed.
- Driver / process crash — systemd
Restart=alwaysor a supervisor that reloads the model and drains the queue. Expect minutes of downtime per crash; the utilization factor already assumes some of this. - Upstream model updates — pin model hashes and ComfyUI node versions. Auto-updating a production worker is how you wake up to a broken graph.
Monitoring (minimum viable)
- Queue depth and age (the real health metric — GPU utilization can read 100% while the queue starves).
- Failure rate by error class, trended.
nvidia-smimemory and temperature; throttle before thermal shutdown.- Disk free on scratch and archive mounts.
- Clips/day and cost/clip computed from your own counters — replace the article's assumptions with measured values within the first week.
Checkpointing
- Workers should checkpoint nothing mid-clip (diffusion steps are not cheap to resume). Instead: checkpoint the queue position — committed jobs stay committed, in-flight jobs requeue cleanly on restart. Store seeds so a requeued job can reproduce or deliberately vary output.
Two configurations
Starter: one RTX 4090, Wan 2.2 14B
| Item | Choice | Why |
|---|---|---|
| GPU | 1× RTX 4090 (24 GB) | Fits Wan 2.2 14B at FP8/GGUF with T5 offload; strongest $/hour rental market |
| Host | 64 GB RAM (for T5 offload), 2 TB NVMe, 850 W+ PSU | T5-XXL offload wants real system RAM; scratch wants fast disk |
| Software | ComfyUI + pinned Wan workflow, Redis queue, systemd supervisor | Standard video graph ecosystem; HTTP API for workers |
| Throughput | ~249 clips/day (75% util assumption) | ~10.5 h finished video/month |
| Cloud cost | ~$248/mo at $0.34/hr | $0.011–0.033/clip |
| Local cost | ~$64/mo electricity + ~$44/mo amortized GPU | ~9-month payback vs renting |
This is the configuration most "automated video channel" projects actually need. Start here, measure for two weeks, then decide whether to scale.
Higher throughput: four workers, or one fat card
| Item | Choice | Trade |
|---|---|---|
| 4× RTX 4090 | 4 independent ComfyUI workers behind one queue | ~4× the starter's throughput (~1,000 clips/day Wan-class); ~$990/mo rented, ~$255/mo electricity + $178/mo amortized owned. Requires a real host (Threadripper/EPYC-class, 128 GB+ RAM, 1600 W PSU or dual PSUs) and serious cooling |
| 1× L40S 48 GB | Single worker, less quantization | Not faster per clip (864 GB/s < 1,008 GB/s) — fewer OOM retries, room for FP16-ish weights and longer clips. ~$577/mo rented |
| 1× RTX 5090 32 GB | Single newest-gen worker | ~1.5–2× 4090 speed in early 480p measurements; 32 GB eases LTX-2 / Wan headroom. ~$438/mo at $0.60/hr |
Multi-GPU on one ComfyUI graph (xDiT-style tensor parallel) exists for HunyuanVideo-class models but adds failure modes; independent single-GPU workers behind a queue are operationally boring and therefore better for unattended runs.
When local beats cloud — decision rule
Steal the breakeven shape from cloud vs self-hosting, with video inputs:
| Situation | Winner |
|---|---|
| Demand is bursty, campaign-driven, or unproven | Cloud — rent for the burst, kill the box after |
| Sustained queue, utilization >~50%, 9+ month horizon | Local — payback under a year, then near-zero marginal GPU cost |
| You lack ops time (updates, crashes, cooling, egress) | Cloud — the premium buys someone else's pager |
| Model needs 48 GB+ and you only generate occasionally | Cloud L40S/A100 — don't buy a card you'll underuse |
| Latency-sensitive interactive preview + batch archive | Hybrid — local 4090 for iteration, cloud burst for archive runs |
Electricity price is a first-class input: at $0.40/kWh (some EU residential rates) local ownership's $64/mo becomes ~$140/mo and payback stretches past 15 months; at $0.10/kWh (some US industrial/rural) it shortens toward 7 months. Re-run the formula with your tariff.
Limitations
- Per-clip times are third-party community measurements (Jul 2025–Jul 2026 reports on RTX 4090/5090/3060), not controlled benchmarks run by OpenGPU Radar. Driver versions, ComfyUI node sets, sampling steps, and CFG all move these numbers 2× in either direction. Measure your own graph before committing capex.
- The 75% utilization factor is our assumption, not measured fleet data. Heavy retry loops, prompt-iteration workloads, or flaky hosts will land lower; perfectly scheduled batches with no failures can land higher.
- Rental rates are dated snapshots (RunPod/CloudRift/computeprices listings, Sep 30 – Oct 2026), not live quotes. Community-cloud rates in particular swing with supply; the $0.34/hr 4090 has been observed as low as $0.16/hr and as high as $0.74/hr (Secure) in the same week.
- VRAM footprints are community-documented ranges for mid-2026 model builds. Offload strategy, attention implementation (sliding tile, sage attention), and VAE tiling change them materially.
- License terms are summarized, not legal advice. HunyuanVideo's territory restrictions and LTX's revenue cap especially — read the current license before shipping commercially.
- No uptime promise is made or implied. Treat clips/day figures as capacity planning inputs under the stated assumptions, not SLAs. Plan for maintenance windows, model migrations, and driver regressions.
- Quality is not modeled. Faster models (LTX) and heavier models (Wan 14B) differ on prompt adherence, motion, and artifact rates; pick by evaluating on your prompt distribution.
- Electricity is US residential average (~18.3¢/kWh, ~Sept 2026, EIA-derived). Swap in your tariff. Cooling load in warm climates adds 10–30% on top of component draw.
Model requirements and generation times compiled from community documentation and measurements (Novita, LocalAI Master, MagicHour, Salad, lilting.ch, RunPod, LTX blog), accessed 2026-10-11. Rental rates from provider listing pages and computeprices.com snapshots dated Sep 30, 2026. Electricity from EIA Electric Power Monthly–derived state tables (US residential average, 2026). Storage rates from Backblaze and Cloudflare published pricing, 2026. Utilization (75%) and system power (600 W load / 100 W idle) are labeled assumptions — replace with measured values.
Related Articles
Cloud GPU vs Self-Hosting: When Renting Actually Wins
Breakeven math for cloud GPU vs owning hardware: hourly cost, utilization, electricity, depreciation, and operational overhead. Worked example serving Llama 3.3 70B on H200.
Cost AnalysisHow Much Does It Cost to Run DeepSeek R1?
DeepSeek R1 costs across API, cloud GPU, and self-hosting. Why the 671B MoE model costs differently than it appears.
Cost AnalysisFLUX.1 [dev] Cost-Per-Image: RTX 4090 vs L40S Pricing Breakdown
FLUX.1 [dev] VRAM sizing, quantization, and cloud spot rates for RTX 4090 vs L40S. Dollar-per-image math with explicit throughput assumptions, not hidden ones.