Model
Trends
.ai
Model
Trends
.ai
the best open source ai models

How much VRAM do you need for AI video? Wan 2.2, HunyuanVideo 1.5 and LTX-2 by GPU size (2026)

Topic: Video models
By P.I. Panda
Published 2026-10-06
Share

Overview

Short answer: 8 GB is enough for short, low-resolution clips with
Wan 2.1
1.3B or a quantised
Wan 2.2 5B
. 24 GB is the real starting point for 720p video with today's best open models. 32 GB runs
Wan 2.2
14B and
LTX-2
at 720p without workarounds. The full-precision versions of the biggest models still need 48 to 80 GB, which means datacenter cards or the cloud.
Video needs so much more memory than images because a single 1024x1024 image is about one million pixels, while a 5-second clip at 24 frames per second is 120 frames: the model works on more than a hundred times as much at once, and has to keep every frame consistent with its neighbours, so it cannot make them one at a time. Three things drive memory and they multiply: model size and precision (full, FP8 or GGUF), resolution (720p has about 2.25 times the pixels of 480p) and clip length (roughly proportional). A model that fits at 480p for 3 seconds can run
out of memory
at 720p for 5 seconds on the same card.
This guide covers the open-weight video models you can download and run yourself. The figures are practical ranges from official model cards and community tests, checked on 6 October 2026; results vary with the
ComfyUI
version, the workflow and how much system RAM is available for offloading. For image models, see the companion guide on
VRAM requirements
by model; both are updated monthly.

Reference

Name
Type
What it is
Wan 2.1
T2V 1.3B
8 GB: 480p, ~5 s clips | 12 GB+: comfortable
The 8 GB option. About 8.19 GB in use, 480p output, around 4 minutes per 5-second clip on an RTX 4090. Quality a clear step below the bigger models. Apache 2.0.
Wan 2.2
TI2V 5B
8 GB: GGUF Q8 with offloading, slow | 12-16 GB: quantised at 720p | 24 GB: full precision (official minimum)
Text or image to 720p at 24 fps. Alibaba lists 24 GB (an RTX 4090) as the minimum at full precision and under 9 minutes for a 5-second 720p clip. The Q8 GGUF is about 5.4 GB; ComfyUI offloads the rest to system RAM. Apache 2.0.
Wan 2.2
A14B (T2V and I2V)
12-16 GB: GGUF Q4-Q5 with offloading, short clips | 24 GB: FP8 at 480p, 720p with the text encoder offloaded | 32 GB: FP8 at 720p | 80 GB: full precision
The best open quality. Mixture of experts: 27B parameters total, 14B active, one expert for the noisy early steps and one for detail; only one expert is in
VRAM
at a time. FP8 is about 14.3 GB per expert, GGUF Q4-Q5 about 10-11 GB. Official requirement 80 GB is for full precision with nothing offloaded. Apache 2.0.
12-16 GB: ~14 GB minimum with offloading | 24 GB: quantised at 720p | 32 GB: comfortable | 28-48 GB: full precision 720p
Tencent, November 2025, 8.3B parameters, 480p or 720p with a built-in
upscaler
to 1080p; 80 GB recommended for best quality. Not the original
HunyuanVideo
(13B), which needs 60-80 GB at full precision; many older guides mix the two up.
LTX-2
(19B) /
LTX-2.3
(22B)
16 GB: draft previews in FP8 (distilled pipeline) | 24 GB: FP4, clips up to ~97 frames | 32 GB: FP8 (official minimum) | 48 GB+: BF16 (~43 GB file)
Lightricks, weights January and March 2026. Video with synced audio, up to 4K, and fast: around 90 seconds for a 5-second clip on an RTX 4090. The trade is memory; Lightricks also recommends 100 GB+ of free disk.
System RAM
32 GB workable minimum | 64 GB comfortable for 14B models on 12-24 GB cards
Offloading moves part of the model to regular RAM, so system memory matters more for video than for images.
Disk
fast NVMe SSD
Video models are tens of GB each; loading from a slow drive adds minutes to every model switch.
Wan
2.5-3.0, Kling, Veo, Runway, Seedance
API only
No downloadable weights, so no VRAM figure; you pay per second of video instead. Wan 2.2 is the newest
Wan
with weights (as of October 2026).
AD
Video without a 24 GB card
BitVector Prism runs the Wan and LTX video models from this guide on its own GPUs: pick the model, upload a frame or type a prompt, get the clip in the browser. Nothing to install, no offloading.

Step by step

  1. Read your VRAM: Windows Task Manager > Performance > GPU > Dedicated GPU memory; nvidia-smi on Linux. Read your system RAM too; 32 GB is the floor for offloading.
  2. Pick the model row for your card from the table: 8 GB means Wan 2.1 1.3B or Wan 2.2 5B GGUF; 12-16 GB adds Wan 2.2 14B GGUF and HunyuanVideo 1.5 with offloading; 24 GB is the common video card (5B full, 14B FP8, LTX-2 FP4); 32 GB runs everything at 720p in FP8.
  3. Download the precision the model page on this site lists for your tier (FP8 for 24-32 GB, GGUF Q4-Q5 for 12-16 GB) rather than the BF16 original.
  4. Start at 480p and short clips (49-81 frames). Get motion and composition right, then upscale or re-render at 720p. Going from 480p to 720p roughly doubles frame memory, so a workflow at 14 GB at 480p may need 24 GB or more at 720p.
  5. For Wan 2.2 14B on 24 GB, move the T5/UMT5 text encoder to system RAM (CPU device on the loader, or ComfyUI's automatic offload). Without it a 720p clip can pass 24 GB partway through.
  6. For anything over 5-6 seconds, generate in segments and chain them: the last frame of one clip becomes the start image of the next. Most open models are tuned for about 5 seconds and lose quality beyond it.
  7. If a clip still does not fit, lower frames before resolution, use tiled VAE decode, and keep the browser and games off the GPU.
  8. If you need Wan 2.2 14B or LTX-2 at 720p regularly and own less than 24 GB, run those on the cloud:
    PirateDiffusion
    runs the Wan,
    LTX
    and H3 models on server GPUs from Telegram;
    BitVector
    from a web app.

Examples

Wan 2.2 14B image-to-video on 24 GB (ComfyUI)

UNETLoader x2: wan2.2_i2v_high_noise_14B_fp8_scaled, wan2.2_i2v_low_noise_14B_fp8_scaled | CLIPLoader: umt5_xxl_fp8_e4m3fn_scaled, device: cpu | Load VAE: wan_2.1_vae
WanImageToVideo: 832x480, 81 frames | KSamplerAdvanced high (steps 0-10, cfg 3.5) -> KSamplerAdvanced low (steps 10-20, cfg 3.5) | shift 5 | VAE Decode (Tiled)

Wan 2.2 5B on an 8-12 GB card

UnetLoaderGGUF: wan2.2_ti2v_5B_Q8_0.gguf | CLIPLoader: umt5_xxl_fp8 (cpu) | 1280x704 is the native 720p shape; start at 832x480, 49 frames, 20 steps
python main.py --lowvram

The same clip on the cloud

/workflow /run:wan22-i2v slow dolly in, she looks up and smiles, soft wind in her hair /length:81 (reply to your image on PirateDiffusion)

Tips

  • Reported times for a single 5-second clip: Wan 2.1 1.3B 480p about 4 minutes on an RTX 4090; Wan 2.2 5B 720p under 9 minutes (official); Wan image-to-video about 12.7 minutes on a 4090 against about 7 on a 5090 (roughly 45 percent faster in one published test); LTX-2 about 90 seconds on a 4090. Offloading can make any of these several times slower.
  • The 80 GB figure Alibaba gives for Wan 2.2 A14B is for full precision with nothing offloaded; FP8 and GGUF plus the text encoder in system RAM is how everyone actually runs it.
  • Two 14B experts means two files: check the model page for which expert a
    LoRA
    was trained on (high-noise or low-noise); loading it on the wrong one does nothing.
  • LTX-2 trades memory for speed and audio; if you can fit FP8 on 32 GB it is the fastest open path to a finished clip with sound.
  • A 5090 launched at 1,999 dollars but sold for roughly 4,500-6,000 in October 2026 and still cannot hold the full-precision models. For occasional video, renting a 48 or 80 GB GPU by the hour, or a flat-rate cloud plan, is cheaper and much faster than offloading on a smaller card.
  • The Wan, Hunyuan and LTX family pages on this site list every downloadable variant with its file size; a variant needs its file size plus 3-4 GB free in VRAM to run at speed.

Troubleshooting

CUDA out of memory partway through a 720p clip on 24 GB

Why it happens
The text encoder is still in VRAM alongside the active 14B expert.

How to fix it
Load the UMT5 encoder on the CPU, or drop to 480p; see the CUDA out of memory guide for the full checklist.

Fits in memory but takes 40 minutes

Why it happens
Heavy offloading to system RAM, or a GGUF Q4 on a card that would run FP8.

How to fix it
Shorten the clip, lower resolution, or use the largest precision that stays resident; check "Shared GPU memory" in Task Manager stays near zero.

HunyuanVideo guide says 60-80 GB

Why it happens
It describes the original 13B HunyuanVideo, not HunyuanVideo 1.5 (8.3B).

How to fix it
Use 1.5: about 14 GB with offloading, 24 GB comfortable.

LTX-2 will not load on 16 GB

Why it happens
Only the distilled FP8 preview pipeline fits; the full model needs 24 GB (FP4) or 32 GB (FP8).

How to fix it
Use the distilled preview for drafts and render finals on a bigger card or the cloud.

Long clip falls apart after 6 seconds

Why it happens
Open models are tuned for about 5 seconds; memory and quality both degrade beyond it.

How to fix it
Chain 5-second segments from the last frame; upscale and interpolate afterwards.

Cannot find Wan 2.5 or 3.0 weights

Why it happens
They are API-only releases.

How to fix it
Wan 2.2 is the newest Wan you can run locally; the Wan family page tracks any change.

AD
PirateDiffusion
Wan 2.2 14B, LTX-2 and H3 from Telegram
PirateDiffusion hosts the large video models at full precision and sends the finished clip to your chat. Unlimited generation on a fixed monthly price, less than a month of 5090 street price per year.

Questions

Can I make AI video with 8 GB of VRAM?

Yes, but only short, low-resolution clips. Wan 2.1 1.3B makes 480p clips in about 8 GB, and Wan 2.2 5B runs as a GGUF file with offloading. Wan 2.2 14B, HunyuanVideo and LTX-2 need more.

Is 24 GB enough for AI video?

For most people, yes. A 24 GB card runs Wan 2.2 5B at full precision, Wan 2.2 14B in FP8 (720p with the text encoder offloaded), HunyuanVideo 1.5 and LTX-2 in FP4. It is the most common card for local video work.

Why does Wan 2.2 say it needs 80 GB?

That is the figure for the 14B model at full precision with nothing offloaded. With FP8 or GGUF files and the text encoder in system RAM it runs on 24 GB, and on 12-16 GB more slowly.

Is an RTX 5090 worth it for AI video?

If you make video often, the extra 8 GB over a 4090 (32 against 24) is what gets you comfortable 720p with Wan 2.2 14B and FP8 LTX-2, and it is roughly 45 percent faster. At two to three times its 1,999 dollar launch price, renting is cheaper for anyone who does not generate video every day.

Does HunyuanVideo need more VRAM than Wan?

HunyuanVideo 1.5 (8.3B) needs less than Wan 2.2 14B: about 14 GB with offloading. The original HunyuanVideo (13B) needs much more, so check which version a guide is talking about.

How much system RAM?

32 GB is the workable minimum for Wan 2.2 and HunyuanVideo with offloading; 64 GB is comfortable for 14B models on a 12-24 GB card.

Links and sources

Models in this guide

Written by
P.I. Panda

Related guides

Errors & fixes
VRAM requirements for every popular AI image and video model (2026): full precision, FP8 and GGUF
How much VRAM you need for SD 1.5, SDXL, SD 3.5, Z-Image Turbo, FLUX.1, FLUX.2 Klein and Dev, Qwen-Image, Wan 2.2, HunyuanVideo 1.5 and LTX-2, at full precision, FP8 and 4-bit GGUF, with the comfortable card and licence for each, what every GPU size from 8 to 48 GB can run, and the order of fixes when a model does not fit.
By P.I. Panda
2026-10-06
Video models
Wan 2.2 video: text-to-video and image-to-video setup, settings and LoRAs
Getting first results from Wan 2.2: which model size to pick (5B or 14B), the high-noise / low-noise expert pair, resolution and frame counts that work, the lightning LoRAs that cut renders to four steps, prompt structure for motion, and how to run it hosted when the card is too small.
By Captain
2026-05-24
Errors & fixes
CUDA out of memory: how much VRAM each AI model needs and how to fit it
A VRAM table for SD 1.5, SDXL, Flux, FLUX.2, Qwen-Image, Z-Image, Wan 2.2 and HunyuanVideo, and the techniques that make a model fit: fp8 and GGUF quantisation, CPU offload, tiled VAE, resolution and frame limits, and when to stop fighting and run it hosted.
By Captain
2026-05-18
Errors & fixes
GPU and VRAM guide for AI image and video: what each model needs and what to buy (or not)
How much graphics memory each model family really needs at usable speed (SD 1.5, SDXL, Flux, FLUX.2, Qwen Image, Z-Image, Wan 2.2, LTX-2, H3), what fp8 and GGUF quantisation buy you, why system RAM and disk matter too, a tier list of cards from 8 to 32 GB, Mac and AMD notes, and the point at which renting is cheaper than buying.
By Quartermaster
2026-01-07
Models
LTX 2.3 Crisp Enhance: sharper, more cinematic video detail
vrgamedevgirl's Crisp Enhance LoRA for LTX 2.3 increases sharpness, micro-detail and contrast in generated video. Learn how to balance it with the Soft Enhance sibling, weights and prompt ideas.
By Captain
2026-10-03
Cloud services
PirateDiffusion: run any model from Telegram, no GPU needed
The commands that matter in the PirateDiffusion Telegram bot: /render with model trigger words, (( )) and [[ ]] weighting, #recipes, /adetailer, /highdef and /facelift upscaling, /remix, /inpaint, ComfyUI workflows with /wf /run:, and the new // skills that pick the model for you.
By Captain
2026-10-03

More guides

Cloud services
Flat fee vs tokens: why unlimited plans like Graydient.ai are the best value for AI creators in 2026
A ranked cost comparison, with prices checked on 8 October 2026, of flat-fee unlimited plans like Graydient.ai against token and credit pricing from Midjourney, Leonardo, fal.ai, Replicate, Runway and Kling: what 1,000 images and 1,000 videos a month really cost on each, and why one fixed fee for unlimited images, video, audio, Grok LLM chat and web apps is the best value for creators who iterate.
By Captain
2026-10-08
Models
FLUX.1 Dev: how to prompt it, best settings and creative ideas
A practical guide to FLUX.1 Dev by Black Forest Labs: natural-language prompting, guidance and step settings, LoRA stacking, text rendering and prompt ideas that play to its strengths.
By Captain
2026-09-30
Models
Stable Diffusion XL 1.0: prompts, settings and what it still does best
How to get the most out of the official SDXL 1.0 base model: native resolutions, CFG and sampler settings, the refiner, negative prompts and ideas for styles where SDXL still shines.
By Captain
2026-10-01
Models
Stable Diffusion 1.5: the classic model, prompted properly
Stable Diffusion 1.5 is still worth knowing: the right resolution, CFG and sampler, how to use its huge library of LoRAs and embeddings, and creative prompt ideas suited to a 512-pixel model.
By Captain
2026-10-02
Models
FLUX.2 Dev: prompting the newest Black Forest Labs model
FLUX.2 Dev brings larger prompts, better text, stronger realism and multi-reference editing. This guide covers the turbo workflow, guidance and resolution settings, prompt structure and ideas that use its new strengths.
By Captain
2026-10-03
Models
Qwen-Image: long prompts, perfect text and bilingual posters
Qwen-Image is the model to reach for when the words in the picture matter. Learn how to write its long descriptive prompts, render English and Chinese text accurately, set CFG and steps, and explore layout-heavy creative ideas.
By Captain
2026-09-29