Model
Trends
.ai
Model
Trends
.ai
the best open source ai models

WanImageToVideo and Wan 2.x video nodes node in ComfyUI

WanImageToVideo builds the conditioning and the empty video latent for a Wan 2.1/2.2 image-to-video render from a start image, a CLIP Vision embedding and the frame count; the surrounding graph uses Load Diffusion Model, Load CLIP, Load VAE and KSampler.
Node pack: ComfyUI core
Category: Video and animation

What WanImageToVideo and Wan 2.x video nodes does

Wan 2.x (Alibaba) is the most used open video model in ComfyUI. The core nodes WanImageToVideo (I2V), WanFunControlToVideo, WanFunInpaintToVideo, WanVaceToVideo (VACE control and reference) and WanFirstLastFrameToVideo take the positive and negative conditioning, the VAE, a width/height/length and the start image (plus a CLIP Vision output from the clip_vision_h encoder) and return conditioning with the image baked in together with a LATENT of the right shape. Text-to-video uses EmptyHunyuanLatentVideo as the latent instead.
The rest is a normal graph: Load Diffusion Model with wan2.1_i2v_480p_14B_fp8 (or the 2.2 high/low noise pair), Load CLIP with umt5_xxl_fp8 (type wan), Load VAE with wan_2.1_vae, CLIP Text Encode, KSampler (uni_pc or euler, 20-30 steps, cfg 5-6, or cfg 1 with a LightX2V/CausVid LoRA at 4-8 steps), VAE Decode and Video Combine at 16 fps. ModelSamplingSD3 with shift 5-8 sits between the model and the sampler.

Inputs

Name
Type
What it is
positive / negative
CONDITIONING
Prompt conditioning from CLIP Text Encode (umt5).
vae
VAE
wan_2.1_vae.safetensors.
clip_vision_output
CLIP_VISION_OUTPUT
From CLIP Vision Encode with clip_vision_h (I2V 2.1).
start_image
IMAGE
The first frame.
width / height / length
INT
Resolution (832x480 or 1280x720) and frames (81 = 5 s at 16 fps; 4n+1).
batch_size
INT
Usually 1.

Outputs

Name
Type
What it is
positive / negative
CONDITIONING
Conditioning with the image reference.
latent
LATENT
Empty video latent of the requested size for KSampler.
AD
Skip the setup: ComfyUI workflows preinstalled on the cloud
BitVector runs ready-made ComfyUI workflows on its own GPUs. No install, no missing nodes, no red boxes. Open it on your phone or laptop and generate in a minute.

How to use WanImageToVideo and Wan 2.x video nodes

  1. Download the Wan 2.1 I2V 14B fp8 diffusion model, umt5_xxl_fp8 text encoder, clip_vision_h and the Wan 2.1 VAE into their folders.
  2. Build loaders: Load Diffusion Model, Load CLIP (type wan), Load VAE, Load CLIP Vision.
  3. Load Image -> CLIP Vision Encode; connect it, the prompts, the VAE and the image to WanImageToVideo (832x480, length 81).
  4. ModelSamplingSD3 shift 8 -> KSampler uni_pc, 20 steps, cfg 5, denoise 1.0.
  5. VAE Decode -> Video Combine at 16 fps, h264.

Settings and tips

  • length must be 4n+1 (17, 33, 49, 65, 81).
  • LightX2V / CausVid distillation LoRAs allow 4-6 steps at cfg 1: ten times faster.
  • Wan 2.2 uses two models (high noise, low noise) with two KSampler (Advanced) nodes split by steps; the 5B TI2V model is a single file for 8-12 GB cards.
  • The 480p model works on 12 GB with fp8 and --lowvram; 720p wants 16-24 GB.
  • Negative prompt matters: the official long Chinese negative (blur, static, distortion) improves motion.

Troubleshooting WanImageToVideo and Wan 2.x video nodes

Error: length must be ... / shape mismatch in the latent

Why it happens
A frame count that is not 4n+1, or width/height not multiples of 16.

How to fix it
Use 81 frames and 832x480 or 1280x720.

Output is a still image with no motion

Why it happens
Shift too low, cfg too low without a distillation LoRA, or a very short prompt.

How to fix it
ModelSamplingSD3 shift 5-8, cfg 5-6 (or a LightX2V LoRA at cfg 1), and describe the motion in the prompt.

clip_vision_output is required / wrong CLIP Vision model

Why it happens
Wan 2.1 I2V needs clip_vision_h.safetensors; SigLIP or ViT-L models from other graphs do not match.

How to fix it
Load clip_vision_h with Load CLIP Vision and encode the start image. Wan 2.2 5B and some 2.2 graphs do not use it at all.

Out of memory at the VAE decode

Why it happens
Decoding 81 frames at 720p in one go.

How to fix it
Use VAE Decode (Tiled) with a temporal tile, lower the resolution, or decode at 480p and upscale frames afterwards.

Load CLIP fails with the umt5 file

Why it happens
The type dropdown on Load CLIP is not set to wan.

How to fix it
Set type to wan (and use the fp8 umt5_xxl file to save memory).

AD
Run this workflow from your phone
Every workflow on this page is preinstalled on BitVector with the models and custom nodes already in place. Pick one, type a prompt, done. Zero setup, nothing to download.

Questions about WanImageToVideo and Wan 2.x video nodes

Which Wan model should I pick?

Wan 2.2 14B I2V (high+low noise, fp8) is the best quality; Wan 2.2 5B TI2V fits small GPUs; Wan 2.1 14B I2V 480p is the stable choice with the most LoRAs.

Can I control the camera or motion?

Yes: Wan Fun Control (depth/pose video), VACE for references and inpainting, and motion LoRAs from the catalog.

Text-to-video?

Use EmptyHunyuanLatentVideo as the latent with the Wan T2V model; the rest of the graph is the same.

Related nodes

Load Diffusion Model
ComfyUI core
Load Diffusion Model (UNETLoader) loads a bare denoising network such as Flux, SD3.5, Wan or HunyuanVideo from models/diffusion_models, with a weight_dtype option for fp8 to save VRAM.
Load VAE
ComfyUI core
Load VAE (VAELoader) loads a standalone VAE file from models/vae so a checkpoint with a missing or weak VAE decodes clean colours, and so split models like Flux and Wan get their decoder.
Video Combine (VideoHelperSuite)
ComfyUI-VideoHelperSuite
Video Combine turns an IMAGE batch of frames into an MP4, WebM, GIF or image sequence at a chosen frame rate, with optional audio, loop count and ping-pong, the standard output node for every video workflow.
KSampler (Advanced)
ComfyUI core
KSampler (Advanced) is KSampler with step-range control: add_noise, start_at_step, end_at_step and return_with_leftover_noise, used for base + refiner chains and two-model sampling.
EmptySD3LatentImage
ComfyUI core
EmptySD3LatentImage creates the 16-channel empty latent that Flux, SD3 and SD3.5 need; it replaces Empty Latent Image in those workflows.
AD
Tired of fixing nodes? Let the cloud do it
BitVector keeps hundreds of ComfyUI workflows installed, updated and tested on fast cloud GPUs. No Python, no CUDA errors, no VRAM limits. Works from any browser.

More ComfyUI nodes

CLIP Text Encode (Prompt)
ComfyUI core
CLIP Text Encode turns a text prompt into CONDITIONING using the model text encoder. One node holds the positive prompt, a second one the negative prompt.
ComfyUI Manager
ComfyUI-Manager
ComfyUI Manager is the extension that installs, updates and fixes custom node packs and models from inside the interface, resolves missing nodes in imported workflows and snapshots your setup.
KSampler
ComfyUI core
KSampler runs the denoising loop: it takes the model, prompts and a latent and produces the finished latent image, controlled by seed, steps, cfg, sampler, scheduler and denoise.
Load Checkpoint
ComfyUI core
Load Checkpoint (CheckpointLoaderSimple) opens a .safetensors or .ckpt model file and hands out the three parts every workflow needs: the diffusion MODEL, the CLIP text encoder and the VAE.
Save Image
ComfyUI core
Save Image writes the IMAGE tensor to ComfyUI/output as a PNG, with the whole workflow embedded in the file metadata so the picture can be dragged back into ComfyUI to restore the graph.
VAE Decode
ComfyUI core
VAE Decode converts the sampled LATENT into a pixel IMAGE with the VAE; VAE Decode (Tiled) does the same in tiles for very large images.