Model
Trends
.ai
Model
Trends
.ai
i migliori modelli di IA open source

Wan 2.2 video: text-to-video and image-to-video setup, settings and LoRAs

Argomento: Modelli video
Di Captain
Pubblicato il 2026-05-24
Getting first results from Wan 2.2: which model size to pick (5B or 14B), the high-noise / low-noise expert pair, resolution and frame counts that work, the lightning LoRAs that cut renders to four steps, prompt structure for motion, and how to run it hosted when the card is too small.
Questa pagina non è ancora tradotta, quindi è mostrata in inglese.

Panoramica

Wan 2.2
is Alibaba's open video family and the most downloaded video model on this site. It comes as a 5B hybrid text-and-image-to-video model built for consumer cards (720p on 8 to 12 GB), and as 14B text-to-video and image-to-video models that use two experts: a high-noise model that lays out motion and composition in the first steps and a low-noise model that adds detail in the last steps. The 14B pair produces the cinematic results people share; the 5B model is the one to learn on.
The community has trained hundreds of
Wan
LoRAs
: camera moves (orbit, dolly, crash zoom), motion styles, characters, effects, and the lightning / distillation LoRAs that reduce 30-step renders to 4 to 8 steps with little loss. The Wan family page ranks them by heat; the lightning LoRAs are nearly always at the top.
Hosted options are a good fit for video because renders are long and VRAM-hungry:
PirateDiffusion
runs Wan workflows from Telegram with /wf /run, and
BitVector
keeps the Wan graphs preinstalled with the LoRAs in place. Both are linked in the Run box of every Wan model page.

Riferimento

Nome
Tipo
Cos'è
Wan2.2-TI2V-5B
model
One model for text- and image-to-video, 720p at 24 fps, 121 frames max. fp16 10 GB, fp8 5 GB. 8 to 12 GB cards.
Wan2.2-T2V-A14B / I2V-A14B
model pair
High-noise and low-noise experts, 14B each (27B total, 14B active). 480p and 720p. fp8 about 14 GB each; GGUF Q4 about 8 GB each.
umt5-xxl text encoder
file
fp8 scaled (6 GB) or fp16 (11 GB). Shared by all Wan models.
Wan 2.1 / 2.2 VAE
file
The 5B model uses the 2.2 VAE; the 14B models use the 2.1 VAE. Mixing them produces garbage.
Resolution
setting
480x832 or 832x480 for fast tests; 720x1280 / 1280x720 for finals. Multiples of 16.
Frames
setting
4n+1: 33, 49, 65, 81. 81 frames at 16 fps is five seconds.
Steps / shift / CFG
setting
20 to 30 steps, shift 5 to 8,
CFG
3.5 to 5 without lightning. With lightning LoRAs: 4 to 8 steps, CFG 1, shift 5.
Lightning / distill LoRA
LoRA
Separate LoRAs for the high-noise and low-noise experts; load each on its own model. Weight 1.0.
Prompt
text
Subject, action, camera, lighting, style, in that order. Describe motion verbs explicitly.
Negative prompt
is used by Wan (unlike
Flux
).
AD
Nessuna installazione: esegui i modelli IA nel cloud
BitVector Prism è il modo più semplice per iniziare: scegli un modello, scrivi un prompt e genera in una web app pulita, senza configurare nulla. BitVector è disponibile anche su Discord e sul web (SpyGlass).

Passo dopo passo

  1. Install the loaders:
    ComfyUI
    core supports Wan 2.2; add ComfyUI-GGUF for quantised files. Update ComfyUI first.
  2. Download for your card: 5B fp8 plus the 2.2 VAE and umt5 fp8 on 8 to 12 GB; the 14B fp8 pair plus the 2.1 VAE on 24 GB; the 14B GGUF Q4 pair on 12 to 16 GB.
  3. Load the official template (Workflow > Browse Templates > Video > Wan 2.2) and replace the loaders with your files. For GGUF swap Load Diffusion Model for Unet Loader (GGUF).
  4. For 14B: two KSampler Advanced nodes. The first runs the high-noise model for steps 0 to N/2 (return with leftover noise), the second runs the low-noise model from N/2 to N. The template wires this already.
  5. First render: 480x832, 33 frames, 20 steps, CFG 4, shift 8, prompt "a corgi runs along a beach at sunset, camera tracks alongside, golden light, shallow depth of field". Expect 2 to 6 minutes.
  6. Add the lightning LoRAs from the Wan family page (one for each expert), set steps to 6 (3 + 3), CFG 1. Renders drop to under a minute at 480p.
  7. Image-to-video: load the I2V pair, feed the picture into WanImageToVideo, keep the prompt about motion and camera rather than describing the picture.
  8. Finals: 720p, 81 frames, interpolate to 32 fps with RIFE or FILM, upscale with a video
    upscaler
    from the
    Upscalers
    family if needed.
  9. Too slow or too big: run the same workflow on BitVector, or on PirateDiffusion with /workflow /show:wan to read the fields and /wf /run:wan your prompt.

Esempi

Text-to-video prompt

A fisherman in a yellow raincoat hauls a net onto a small wooden boat in heavy rain, waves rock the boat, the camera is handheld and close, grey stormy light, cinematic, 35mm film grain
Negative: overexposed, static, blurry details, subtitles, worst quality, low quality, deformed, extra fingers, still image

Image-to-video prompt (motion only)

The woman turns her head toward the camera and smiles, wind moves her hair, slow push-in, soft daylight

PirateDiffusion

/workflow /show:wan-i2v
/wf /run:wan-i2v /initimage:Iabc123 the woman turns her head toward the camera and smiles, wind moves her hair, slow push-in

Consigli

  • Describe motion with verbs and the camera with film words (slow dolly in, handheld, static wide shot). Wan follows camera language unusually well.
  • Keep the negative prompt: the official one lists overexposure, static, blurry details, subtitles, worst quality, extra fingers, and it helps.
  • For I2V, generate the first frame with Flux,
    Qwen
    or
    Z-Image
    at the target aspect ratio; a sharp, well-lit still produces a better clip than any prompt.
  • Frame count costs more than resolution: 81 frames at 480p is cheaper than 49 frames at 720p.
  • Motion LoRAs for Wan are tagged by the expert they target; a low-noise-only LoRA adds detail, a high-noise LoRA changes motion.
  • Seed matters a lot in video; keep the seed fixed while you tune the prompt, then vary it for takes.
  • Save the workflow with your LoRA set as a template; Wan graphs are long and easy to break.

Risoluzione dei problemi

Out of memory at the first sampler

Perché succede
fp16 14B weights on a 24 GB or smaller card.

Come risolverlo
fp8 or GGUF Q4 pair, fewer frames, 480p, --lowvram; or the 5B model.

Video is a blurry smear or colour noise

Perché succede
Wrong VAE (2.2 VAE with 14B or 2.1 VAE with 5B), or lightning LoRA without CFG 1.

Come risolverlo
Match the VAE to the model; with lightning LoRAs set CFG 1 and 4 to 8 steps.

Almost no motion

Perché succede
Prompt describes a scene, not an action; or too few steps on the high-noise expert.

Come risolverlo
Add motion verbs and camera moves; give the high-noise expert half the steps.

Subject morphs halfway through

Perché succede
Too many frames for the resolution, or two conflicting LoRAs.

Come risolverlo
Shorten to 49 frames, one motion LoRA at a time, raise steps on the low-noise expert.

Decode fails after sampling succeeded

Perché succede
VAE decode is the memory peak for video.

Come risolverlo
Wan VAE decode in tiles (the tiled option on the decode node), or fewer frames.

LoRA key not loaded warnings

Perché succede
A 2.1 LoRA on 2.2, or a 14B LoRA on the 5B model.

Come risolverlo
Match version and size; the model page names both.

AD
PirateDiffusion
Nessuna installazione: esegui i modelli IA nel cloud
PirateDiffusion è solo su Telegram ed è pensato per i professionisti: migliaia di modelli, LoRA e workflow tramite comandi in chat, con generazione illimitata a prezzo fisso.

Domande

5B or 14B?

5B to learn and for 8 to 12 GB cards; 14B for quality when you have 24 GB or run hosted. Both take the same prompts.

How long is a clip?

Up to 81 frames (5 s at 16 fps) for 14B and 121 frames (5 s at 24 fps) for 5B per render; longer videos are stitched from last-frame continuations.

Does Wan do audio?

Wan 2.2 S2V animates a speaker from audio but does not generate sound;
LTX-2
and
MiniMax H3
generate audio with the video (their guides are on this site).

Where do the lightning LoRAs come from?

Distillation releases by Lightx2v and others; the Wan family page lists them with heat scores and the expert each one targets.

Can I run Wan without a GPU?

Yes: PirateDiffusion from Telegram, BitVector in the browser. Both keep the workflows and LoRAs installed.

Link e fonti

Modelli di questa guida

Scritto da
Captain

Guide correlate

Errori e soluzioni
CUDA out of memory: how much VRAM each AI model needs and how to fit it
A VRAM table for SD 1.5, SDXL, Flux, FLUX.2, Qwen-Image, Z-Image, Wan 2.2 and HunyuanVideo, and the techniques that make a model fit: fp8 and GGUF quantisation, CPU offload, tiled VAE, resolution and frame limits, and when to stop fighting and run it hosted.
Di Captain
2026-05-18
Errori e soluzioni
ComfyUI errors and how to fix them: red nodes, missing models, CUDA out of memory
The ComfyUI error messages people hit first, what each one means and the fix that works: missing custom nodes (red boxes), "Prompt outputs failed validation", CUDA out of memory, wrong model type in a loader, mat1 and mat2 shape errors, header deserialization, torch and xformers mismatches.
Di Captain
2026-05-12
Modelli video
MiniMax H3 video prompts: the official structure
How MiniMax wants an H3 prompt written: the alignment line for image-anchored videos, the three core fields, shots and cuts, camera moves, speaker IDs and dialogue tags, the soundscape and the music. Condensed from the official prompt writing guide.
Di Captain
2026-09-14
Modelli
Flux vs SDXL vs SD 1.5 vs Qwen vs Z-Image: which model family to use in 2026
A practical comparison of the open image model families: prompt following, realism, anime and illustration, text rendering, speed, VRAM, LoRA ecosystem and licence, with a recommendation for each kind of work and hardware.
Di Captain
2026-05-22

Altre guide

Modelli
FLUX.1 Dev: how to prompt it, best settings and creative ideas
A practical guide to FLUX.1 Dev by Black Forest Labs: natural-language prompting, guidance and step settings, LoRA stacking, text rendering and prompt ideas that play to its strengths.
Di Captain
2026-09-30
Modelli
Stable Diffusion XL 1.0: prompts, settings and what it still does best
How to get the most out of the official SDXL 1.0 base model: native resolutions, CFG and sampler settings, the refiner, negative prompts and ideas for styles where SDXL still shines.
Di Captain
2026-10-01
Modelli
Stable Diffusion 1.5: the classic model, prompted properly
Stable Diffusion 1.5 is still worth knowing: the right resolution, CFG and sampler, how to use its huge library of LoRAs and embeddings, and creative prompt ideas suited to a 512-pixel model.
Di Captain
2026-10-02
Modelli
FLUX.2 Dev: prompting the newest Black Forest Labs model
FLUX.2 Dev brings larger prompts, better text, stronger realism and multi-reference editing. This guide covers the turbo workflow, guidance and resolution settings, prompt structure and ideas that use its new strengths.
Di Captain
2026-10-03
Modelli
Qwen-Image: long prompts, perfect text and bilingual posters
Qwen-Image is the model to reach for when the words in the picture matter. Learn how to write its long descriptive prompts, render English and Chinese text accurately, set CFG and steps, and explore layout-heavy creative ideas.
Di Captain
2026-09-29
Modelli
SDXL-Lightning: four-step generation on any SDXL checkpoint
ByteDance's SDXL-Lightning LoRA turns a 30-step SDXL render into a 4-step one. Learn the step counts, the CFG you must use, the sampler, and how to combine it with your favorite checkpoints and style LoRAs.
Di Captain
2026-09-30