WanImageToVideo and Wan 2.x video nodes-node in ComfyUI
WanImageToVideo builds the conditioning and the empty video latent for a Wan 2.1/2.2 image-to-video render from a start image, a CLIP Vision embedding and the frame count; the surrounding graph uses Load Diffusion Model, Load CLIP, Load VAE and KSampler.
Nodepakket: ComfyUI core
Categorie: Video en animatie
Deze pagina is nog niet vertaald en wordt daarom in het Engels getoond.
Wat WanImageToVideo and Wan 2.x video nodes doet
Wan 2.x (Alibaba) is the most used open video model in ComfyUI. The core nodes WanImageToVideo (I2V), WanFunControlToVideo, WanFunInpaintToVideo, WanVaceToVideo (VACE control and reference) and WanFirstLastFrameToVideo take the positive and negative conditioning, the VAE, a width/height/length and the start image (plus a CLIP Vision output from the clip_vision_h encoder) and return conditioning with the image baked in together with a LATENT of the right shape. Text-to-video uses EmptyHunyuanLatentVideo as the latent instead.
The rest is a normal graph: Load Diffusion Model with wan2.1_i2v_480p_14B_fp8 (or the 2.2 high/low noise pair), Load CLIP with umt5_xxl_fp8 (type wan), Load VAE with wan_2.1_vae, CLIP Text Encode, KSampler (uni_pc or euler, 20-30 steps, cfg 5-6, or cfg 1 with a LightX2V/CausVid LoRA at 4-8 steps), VAE Decode and Video Combine at 16 fps. ModelSamplingSD3 with shift 5-8 sits between the model and the sampler.
Invoer
Naam | Type | Wat het is |
|---|---|---|
positive / negative | CONDITIONING | Prompt conditioning from CLIP Text Encode (umt5). |
vae | VAE | wan_2.1_vae.safetensors. |
clip_vision_output | CLIP_VISION_OUTPUT | From CLIP Vision Encode with clip_vision_h (I2V 2.1). |
start_image | IMAGE | The first frame. |
width / height / length | INT | Resolution (832x480 or 1280x720) and frames (81 = 5 s at 16 fps; 4n+1). |
batch_size | INT | Usually 1. |
Uitvoer
Naam | Type | Wat het is |
|---|---|---|
positive / negative | CONDITIONING | Conditioning with the image reference. |
latent | LATENT | Empty video latent of the requested size for KSampler. |
AD

Sla de installatie over: ComfyUI-workflows voorgeïnstalleerd in de cloud
BitVector draait kant-en-klare ComfyUI-workflows op eigen GPU's. Geen installatie, geen ontbrekende nodes, geen rode vakken. Open het op je telefoon of laptop en genereer binnen een minuut.
Zo gebruik je WanImageToVideo and Wan 2.x video nodes
- Download the Wan 2.1 I2V 14B fp8 diffusion model, umt5_xxl_fp8 text encoder, clip_vision_h and the Wan 2.1 VAE into their folders.
- Build loaders: Load Diffusion Model, Load CLIP (type wan), Load VAE, Load CLIP Vision.
- Load Image -> CLIP Vision Encode; connect it, the prompts, the VAE and the image to WanImageToVideo (832x480, length 81).
- ModelSamplingSD3 shift 8 -> KSampler uni_pc, 20 steps, cfg 5, denoise 1.0.
- VAE Decode -> Video Combine at 16 fps, h264.
Instellingen en tips
- length must be 4n+1 (17, 33, 49, 65, 81).
- LightX2V / CausVid distillation LoRAs allow 4-6 steps at cfg 1: ten times faster.
- Wan 2.2 uses two models (high noise, low noise) with two KSampler (Advanced) nodes split by steps; the 5B TI2V model is a single file for 8-12 GB cards.
- The 480p model works on 12 GB with fp8 and --lowvram; 720p wants 16-24 GB.
- Negative prompt matters: the official long Chinese negative (blur, static, distortion) improves motion.
Problemen met WanImageToVideo and Wan 2.x video nodes oplossen
Error: length must be ... / shape mismatch in the latent
Waarom het gebeurt
A frame count that is not 4n+1, or width/height not multiples of 16.
Zo los je het op
Use 81 frames and 832x480 or 1280x720.
Output is a still image with no motion
Waarom het gebeurt
Shift too low, cfg too low without a distillation LoRA, or a very short prompt.
Zo los je het op
ModelSamplingSD3 shift 5-8, cfg 5-6 (or a LightX2V LoRA at cfg 1), and describe the motion in the prompt.
clip_vision_output is required / wrong CLIP Vision model
Waarom het gebeurt
Wan 2.1 I2V needs clip_vision_h.safetensors; SigLIP or ViT-L models from other graphs do not match.
Zo los je het op
Load clip_vision_h with Load CLIP Vision and encode the start image. Wan 2.2 5B and some 2.2 graphs do not use it at all.
Out of memory at the VAE decode
Waarom het gebeurt
Decoding 81 frames at 720p in one go.
Zo los je het op
Use VAE Decode (Tiled) with a temporal tile, lower the resolution, or decode at 480p and upscale frames afterwards.
Load CLIP fails with the umt5 file
Waarom het gebeurt
The type dropdown on Load CLIP is not set to wan.
Zo los je het op
Set type to wan (and use the fp8 umt5_xxl file to save memory).
AD

Start deze workflow vanaf je telefoon
Elke workflow op deze pagina staat voorgeïnstalleerd op BitVector, met modellen en custom nodes al op hun plek. Kies er een, typ een prompt, klaar. Nul installatie, niets te downloaden.
Vragen over WanImageToVideo and Wan 2.x video nodes
Which Wan model should I pick?
Wan 2.2 14B I2V (high+low noise, fp8) is the best quality; Wan 2.2 5B TI2V fits small GPUs; Wan 2.1 14B I2V 480p is the stable choice with the most LoRAs.
Can I control the camera or motion?
Yes: Wan Fun Control (depth/pose video), VACE for references and inpainting, and motion LoRAs from the catalog.
Text-to-video?
Use EmptyHunyuanLatentVideo as the latent with the Wan T2V model; the rest of the graph is the same.
Verwante nodes
Load Diffusion Model
ComfyUI core
Load Diffusion Model (UNETLoader) loads a bare denoising network such as Flux, SD3.5, Wan or HunyuanVideo from models/diffusion_models, with a weight_dtype option for fp8 to save VRAM.
Load VAE
ComfyUI core
Load VAE (VAELoader) loads a standalone VAE file from models/vae so a checkpoint with a missing or weak VAE decodes clean colours, and so split models like Flux and Wan get their decoder.
Video Combine (VideoHelperSuite)
ComfyUI-VideoHelperSuite
Video Combine turns an IMAGE batch of frames into an MP4, WebM, GIF or image sequence at a chosen frame rate, with optional audio, loop count and ping-pong, the standard output node for every video workflow.
KSampler (Advanced)
ComfyUI core
KSampler (Advanced) is KSampler with step-range control: add_noise, start_at_step, end_at_step and return_with_leftover_noise, used for base + refiner chains and two-model sampling.
EmptySD3LatentImage
ComfyUI core
EmptySD3LatentImage creates the 16-channel empty latent that Flux, SD3 and SD3.5 need; it replaces Empty Latent Image in those workflows.
AD

Nodes repareren beu? Laat de cloud het doen
BitVector houdt honderden ComfyUI-workflows geïnstalleerd, bijgewerkt en getest op snelle cloud-GPU's. Geen Python, geen CUDA-fouten, geen VRAM-limieten. Werkt in elke browser.
Meer ComfyUI-nodes
CLIP Text Encode (Prompt)
ComfyUI core
CLIP Text Encode turns a text prompt into CONDITIONING using the model text encoder. One node holds the positive prompt, a second one the negative prompt.
ComfyUI Manager
ComfyUI-Manager
ComfyUI Manager is the extension that installs, updates and fixes custom node packs and models from inside the interface, resolves missing nodes in imported workflows and snapshots your setup.
KSampler
ComfyUI core
KSampler runs the denoising loop: it takes the model, prompts and a latent and produces the finished latent image, controlled by seed, steps, cfg, sampler, scheduler and denoise.
Load Checkpoint
ComfyUI core
Load Checkpoint (CheckpointLoaderSimple) opens a .safetensors or .ckpt model file and hands out the three parts every workflow needs: the diffusion MODEL, the CLIP text encoder and the VAE.
Save Image
ComfyUI core
Save Image writes the IMAGE tensor to ComfyUI/output as a PNG, with the whole workflow embedded in the file metadata so the picture can be dragged back into ComfyUI to restore the graph.
VAE Decode
ComfyUI core
VAE Decode converts the sampled LATENT into a pixel IMAGE with the VAE; VAE Decode (Tiled) does the same in tiles for very large images.

Model
Trends
.ai
© ModelTrends.ai
|
Gemaakt in Japan
|
© 2026