Wan 2.2 video: text-to-video and image-to-video setup, settings and LoRAs
Ämne: Videomodeller
Av Captain
Publicerad 2026-05-24
Getting first results from Wan 2.2: which model size to pick (5B or 14B), the high-noise / low-noise expert pair, resolution and frame counts that work, the lightning LoRAs that cut renders to four steps, prompt structure for motion, and how to run it hosted when the card is too small.
Den här sidan har inte översatts ännu, så den visas på engelska.
Översikt
Wan 2.2
is Alibaba's open video family and the most downloaded video model on this site. It comes as a 5B hybrid text-and-image-to-video model built for consumer cards (720p on 8 to 12 GB), and as 14B text-to-video and image-to-video models that use two experts: a high-noise model that lays out motion and composition in the first steps and a low-noise model that adds detail in the last steps. The 14B pair produces the cinematic results people share; the 5B model is the one to learn on.
The community has trained hundreds of
Wan
LoRAs
: camera moves (orbit, dolly, crash zoom), motion styles, characters, effects, and the lightning / distillation LoRAs that reduce 30-step renders to 4 to 8 steps with little loss. The Wan family page ranks them by heat; the lightning LoRAs are nearly always at the top.
Hosted options are a good fit for video because renders are long and VRAM-hungry:
PirateDiffusion
runs Wan workflows from Telegram with /wf /run, and
BitVector
keeps the Wan graphs preinstalled with the LoRAs in place. Both are linked in the Run box of every Wan model page.
Referens
Namn | Typ | Vad det är |
|---|---|---|
Wan2.2-TI2V-5B | model | One model for text- and image-to-video, 720p at 24 fps, 121 frames max. fp16 10 GB, fp8 5 GB. 8 to 12 GB cards. |
Wan2.2-T2V-A14B / I2V-A14B | model pair | High-noise and low-noise experts, 14B each (27B total, 14B active). 480p and 720p. fp8 about 14 GB each; GGUF Q4 about 8 GB each. |
umt5-xxl text encoder | file | fp8 scaled (6 GB) or fp16 (11 GB). Shared by all Wan models. |
Wan 2.1 / 2.2 VAE | file | The 5B model uses the 2.2 VAE; the 14B models use the 2.1 VAE. Mixing them produces garbage. |
Resolution | setting | 480x832 or 832x480 for fast tests; 720x1280 / 1280x720 for finals. Multiples of 16. |
Frames | setting | 4n+1: 33, 49, 65, 81. 81 frames at 16 fps is five seconds. |
Steps / shift / CFG | setting | 20 to 30 steps, shift 5 to 8, CFG 3.5 to 5 without lightning. With lightning LoRAs: 4 to 8 steps, CFG 1, shift 5. |
Lightning / distill LoRA | LoRA | Separate LoRAs for the high-noise and low-noise experts; load each on its own model. Weight 1.0. |
Prompt | text | Subject, action, camera, lighting, style, in that order. Describe motion verbs explicitly. Negative prompt is used by Wan (unlike Flux ). |
AD

Inget att installera - kör AI-modeller i molnet
BitVector Prism är det enklaste sättet att komma igång: välj en modell, skriv en prompt och generera i en ren webbapp, utan någon installation alls. BitVector finns också på Discord och på webben (SpyGlass).
Steg för steg
- Install the loaders:ComfyUIcore supports Wan 2.2; add ComfyUI-GGUF for quantised files. Update ComfyUI first.
- Download for your card: 5B fp8 plus the 2.2 VAE and umt5 fp8 on 8 to 12 GB; the 14B fp8 pair plus the 2.1 VAE on 24 GB; the 14B GGUF Q4 pair on 12 to 16 GB.
- Load the official template (Workflow > Browse Templates > Video > Wan 2.2) and replace the loaders with your files. For GGUF swap Load Diffusion Model for Unet Loader (GGUF).
- For 14B: two KSampler Advanced nodes. The first runs the high-noise model for steps 0 to N/2 (return with leftover noise), the second runs the low-noise model from N/2 to N. The template wires this already.
- First render: 480x832, 33 frames, 20 steps, CFG 4, shift 8, prompt "a corgi runs along a beach at sunset, camera tracks alongside, golden light, shallow depth of field". Expect 2 to 6 minutes.
- Add the lightning LoRAs from the Wan family page (one for each expert), set steps to 6 (3 + 3), CFG 1. Renders drop to under a minute at 480p.
- Image-to-video: load the I2V pair, feed the picture into WanImageToVideo, keep the prompt about motion and camera rather than describing the picture.
- Too slow or too big: run the same workflow on BitVector, or on PirateDiffusion with /workflow /show:wan to read the fields and /wf /run:wan your prompt.
Exempel
Text-to-video prompt
A fisherman in a yellow raincoat hauls a net onto a small wooden boat in heavy rain, waves rock the boat, the camera is handheld and close, grey stormy light, cinematic, 35mm film grain
Negative: overexposed, static, blurry details, subtitles, worst quality, low quality, deformed, extra fingers, still image
Image-to-video prompt (motion only)
The woman turns her head toward the camera and smiles, wind moves her hair, slow push-in, soft daylight
PirateDiffusion
/workflow /show:wan-i2v
/wf /run:wan-i2v /initimage:Iabc123 the woman turns her head toward the camera and smiles, wind moves her hair, slow push-in
Tips
- Describe motion with verbs and the camera with film words (slow dolly in, handheld, static wide shot). Wan follows camera language unusually well.
- Keep the negative prompt: the official one lists overexposure, static, blurry details, subtitles, worst quality, extra fingers, and it helps.
- Frame count costs more than resolution: 81 frames at 480p is cheaper than 49 frames at 720p.
- Motion LoRAs for Wan are tagged by the expert they target; a low-noise-only LoRA adds detail, a high-noise LoRA changes motion.
- Seed matters a lot in video; keep the seed fixed while you tune the prompt, then vary it for takes.
- Save the workflow with your LoRA set as a template; Wan graphs are long and easy to break.
Felsökning
Out of memory at the first sampler
Varför det händer
fp16 14B weights on a 24 GB or smaller card.
Så löser du det
fp8 or GGUF Q4 pair, fewer frames, 480p, --lowvram; or the 5B model.
Video is a blurry smear or colour noise
Varför det händer
Wrong VAE (2.2 VAE with 14B or 2.1 VAE with 5B), or lightning LoRA without CFG 1.
Så löser du det
Match the VAE to the model; with lightning LoRAs set CFG 1 and 4 to 8 steps.
Almost no motion
Varför det händer
Prompt describes a scene, not an action; or too few steps on the high-noise expert.
Så löser du det
Add motion verbs and camera moves; give the high-noise expert half the steps.
Subject morphs halfway through
Varför det händer
Too many frames for the resolution, or two conflicting LoRAs.
Så löser du det
Shorten to 49 frames, one motion LoRA at a time, raise steps on the low-noise expert.
Decode fails after sampling succeeded
Varför det händer
VAE decode is the memory peak for video.
Så löser du det
Wan VAE decode in tiles (the tiled option on the decode node), or fewer frames.
LoRA key not loaded warnings
Varför det händer
A 2.1 LoRA on 2.2, or a 14B LoRA on the 5B model.
Så löser du det
Match version and size; the model page names both.
AD

Inget att installera - kör AI-modeller i molnet
PirateDiffusion finns bara på Telegram och är byggt för proffs: tusentals modeller, LoRA och arbetsflöden som styrs med chattkommandon, med obegränsad generering till ett fast pris.
Frågor
5B or 14B?
5B to learn and for 8 to 12 GB cards; 14B for quality when you have 24 GB or run hosted. Both take the same prompts.
How long is a clip?
Up to 81 frames (5 s at 16 fps) for 14B and 121 frames (5 s at 24 fps) for 5B per render; longer videos are stitched from last-frame continuations.
Does Wan do audio?
Wan 2.2 S2V animates a speaker from audio but does not generate sound;
LTX-2
and
MiniMax H3
generate audio with the video (their guides are on this site).
Where do the lightning LoRAs come from?
Distillation releases by Lightx2v and others; the Wan family page lists them with heat scores and the expert each one targets.
Can I run Wan without a GPU?
Yes: PirateDiffusion from Telegram, BitVector in the browser. Both keep the workflows and LoRAs installed.
Länkar och källor
- Wan 2.2 on GitHub
- Wan 2.2 weights on Hugging Face
- ComfyUI Wan 2.2 examples
- Wan family page
- Hosted Wan workflows on PirateDiffusion
- BitVector
Modeller i den här guiden
Skriven av
Captain
Relaterade guider
Fel och lösningar
CUDA out of memory: how much VRAM each AI model needs and how to fit it
A VRAM table for SD 1.5, SDXL, Flux, FLUX.2, Qwen-Image, Z-Image, Wan 2.2 and HunyuanVideo, and the techniques that make a model fit: fp8 and GGUF quantisation, CPU offload, tiled VAE, resolution and frame limits, and when to stop fighting and run it hosted.
Av Captain
2026-05-18
Fel och lösningar
ComfyUI errors and how to fix them: red nodes, missing models, CUDA out of memory
The ComfyUI error messages people hit first, what each one means and the fix that works: missing custom nodes (red boxes), "Prompt outputs failed validation", CUDA out of memory, wrong model type in a loader, mat1 and mat2 shape errors, header deserialization, torch and xformers mismatches.
Av Captain
2026-05-12
Videomodeller
MiniMax H3 video prompts: the official structure
How MiniMax wants an H3 prompt written: the alignment line for image-anchored videos, the three core fields, shots and cuts, camera moves, speaker IDs and dialogue tags, the soundscape and the music. Condensed from the official prompt writing guide.
Av Captain
2026-09-14
Modeller
Flux vs SDXL vs SD 1.5 vs Qwen vs Z-Image: which model family to use in 2026
A practical comparison of the open image model families: prompt following, realism, anime and illustration, text rendering, speed, VRAM, LoRA ecosystem and licence, with a recommendation for each kind of work and hardware.
Av Captain
2026-05-22
Fler guider
Modeller
FLUX.1 Dev: how to prompt it, best settings and creative ideas
A practical guide to FLUX.1 Dev by Black Forest Labs: natural-language prompting, guidance and step settings, LoRA stacking, text rendering and prompt ideas that play to its strengths.
Av Captain
2026-09-30
Modeller
Stable Diffusion XL 1.0: prompts, settings and what it still does best
How to get the most out of the official SDXL 1.0 base model: native resolutions, CFG and sampler settings, the refiner, negative prompts and ideas for styles where SDXL still shines.
Av Captain
2026-10-01
Modeller
Stable Diffusion 1.5: the classic model, prompted properly
Stable Diffusion 1.5 is still worth knowing: the right resolution, CFG and sampler, how to use its huge library of LoRAs and embeddings, and creative prompt ideas suited to a 512-pixel model.
Av Captain
2026-10-02
Modeller
FLUX.2 Dev: prompting the newest Black Forest Labs model
FLUX.2 Dev brings larger prompts, better text, stronger realism and multi-reference editing. This guide covers the turbo workflow, guidance and resolution settings, prompt structure and ideas that use its new strengths.
Av Captain
2026-10-03
Modeller
Qwen-Image: long prompts, perfect text and bilingual posters
Qwen-Image is the model to reach for when the words in the picture matter. Learn how to write its long descriptive prompts, render English and Chinese text accurately, set CFG and steps, and explore layout-heavy creative ideas.
Av Captain
2026-09-29
Modeller
SDXL-Lightning: four-step generation on any SDXL checkpoint
ByteDance's SDXL-Lightning LoRA turns a 30-step SDXL render into a 4-step one. Learn the step counts, the CFG you must use, the sampler, and how to combine it with your favorite checkpoints and style LoRAs.
Av Captain
2026-09-30

Model
Trends
.ai
© ModelTrends.ai
|
Tillverkad i Japan
|
© 2026