Model
Trends
.ai
Model
Trends
.ai
les meilleurs modèles d'IA open source

AI video generation for beginners: image-to-video, text-to-video and the first clip with Wan, LTX-2 and H3

Sujet : Modèles vidéo
Par Lookout
Publié le 2026-04-29
A first week with open video models: the difference between text-to-video and image-to-video and why to start with the latter, the three model families worth learning (Wan 2.2, LTX-2, MiniMax H3), frame counts and resolutions that work, how to write a motion prompt, what to expect from a 16 or 24 GB card against the cloud, and how to turn five-second clips into something longer.
Cette page n'est pas encore traduite ; elle est affichée en anglais.

Aperçu

Video models are image models with a time axis: the same denoising loop runs on a block of frames at once, so everything you know about prompts, seeds and denoise carries over, and so do the costs, multiplied by the frame count. The practical consequences for a beginner are three. Start with image-to-video (I2V), because a good still from an image model fixes composition and style and leaves the video model only the motion to invent. Keep clips short (about five seconds), because that is what the models were trained on. And expect to run the large models on the cloud unless you own a 24 GB card.
Three open families cover most needs in 2026.
Wan 2.2
(Alibaba) is the all-rounder: a 5B model that runs on 12-16 GB and a 14B two-expert model with the best open quality, plus a huge
LoRA
ecosystem for motion and style.
LTX-2
(Lightricks) is fast and generates synced audio but wants 24-32 GB.
MiniMax H3
is the audio-video model with dialogue and sound effects, mostly used through hosted workflows; the H3 prompting guides on this site by Mark White go deep on it.
Hunyuan
Video 1.5 is a lighter alternative to
Wan
14B.
The catalog lists each video model and its LoRAs with the base and size they target (5B or 14B, high-noise or low-noise expert), and the model pages show example clips with their prompts. The Run button opens the model on
PirateDiffusion
, where a workflow command produces the clip from a reply to your image, or on
BitVector
.

Référence

Nom
Type
Rôle
I2V (image to video)
mode
Animate a still. Best first mode: composition is already right; the prompt describes motion only.
T2V (text to video)
mode
From a prompt alone. More variety, less control; use for ideas and B-roll.
FLF2V / first-last frame
mode
Two keyframes; the model fills the motion between. Controlled transitions.
Frames and fps
setting
Wan: 81 frames at 16 fps (5 s); LTX-2: 97-121 frames at 24-25 fps; H3: 145 frames at ~24 fps. Memory scales with frames.
Resolution
setting
Wan 2.2: 832x480 or 480x832 (fast), 1280x720 (24 GB+); LTX-2 up to 1080p+ with 32 GB. Portrait for people, landscape for scenes.
Motion prompt
prompt
One subject motion plus one camera motion: "she turns toward the window and smiles; slow push-in". Avoid listing scene details already in the image.
Steps / CFG / shift
setting
Wan 2.2 14B: 20 steps split 10/10 across experts,
CFG
3.5, shift 5; 5B: 20-30 steps, CFG 5; LTX-2: 25-30 steps with its own schedule; distilled LoRAs (lightx2v) bring Wan to 4-8 steps at CFG 1.
Motion LoRAs
file
Camera moves, dances, gestures, physics; matched to 5B/14B and to the expert. Pages here name the target.
Interpolation and upscale
post
RIFE or FILM to 32-60 fps; frame-wise ESRGAN or SeedVR for resolution.
AD
Your first clip without a GPU
BitVector Prism animates a still with the Wan models from this guide in the browser: upload the image, write the motion sentence, download the clip. Nothing to install, works from a laptop.

Pas à pas

  1. Make or pick a still: a 1024x1024 or 832x1216 render with a clear subject and simple background. Faces larger than a tenth of the frame animate best.
  2. Open a Wan 2.2 I2V model page on this site, copy the example motion prompt pattern, and press the Run button (PirateDiffusion: reply to your image with the /workflow command; BitVector: upload and pick the model).
  3. Write a motion prompt of one sentence for the subject and one for the camera. Nothing about clothes, colours or lighting; the image has those.
  4. Generate at 832x480 (landscape) or 480x832 (portrait), 81 frames, default steps. Watch the whole clip; judge motion, not sharpness.
  5. Iterate on the motion prompt only: slower verbs ("slowly", "gently") reduce flailing; a single camera instruction prevents jumps.
  6. Locally, with a 16 GB card: Wan 2.2 5B fp8 or GGUF, 832x480, 49-81 frames, text encoder on CPU. With 24 GB: 14B fp8 at 480p, 720p with the encoder offloaded. See the
    VRAM for AI video
    guide.
  7. Longer pieces: take the last frame of clip one as the start image of clip two with a continuing motion prompt; join in an editor; interpolate to 30 fps and upscale at the end.
  8. Add sound: LTX-2 and H3 generate audio with the clip; for Wan, add music or effects in the editor.

Exemples

Motion prompt for I2V

Positive: she lifts the coffee cup, takes a sip and looks out of the window; slow push-in, shallow depth of field, natural motion
Negative (Wan): static, blurry, distorted face, extra fingers, jump cut, text

PirateDiffusion workflow command

/workflow /run:wan22-i2v she lifts the coffee cup and looks out of the window, slow push-in /length:81 (as a reply to the still)

ComfyUI Wan 2.2 5B on 16 GB

UnetLoaderGGUF wan2.2_ti2v_5B_Q8_0 | CLIPLoader umt5_xxl_fp8 (cpu) | Wan VAE 2.2 | WanImageToVideo 832x480, 81 frames | KSampler 24 steps, cfg 5, uni_pc, simple | VAE Decode (Tiled)

Astuces

  • Motion verbs matter more than adjectives: "turns, walks, lifts, pours, waves" each produce distinct, reliable motion; "beautiful cinematic" produces nothing.
  • One camera move per clip: push-in, pull-back, pan left, orbit. Two moves give a wobble.
  • Portrait stills animate faces better; landscape stills animate environments better.
  • Wan LoRAs for the high-noise expert shape the motion; low-noise ones shape detail. Many packs ship both; load each on its expert.
  • Distilled LoRAs (lightx2v, FastWan) cut Wan from 20 steps to 4-6 at CFG 1 with little quality loss; the model pages list them.
  • Expect 3-10 minutes per clip on a 4090 for Wan 14B and about 90 seconds for LTX-2; the cloud returns most clips in a minute or two.

Dépannage

The clip is a still with a slight shimmer

Pourquoi cela arrive
No motion verb in the prompt, or a static LoRA weight too high.

Comment le corriger
Add a subject motion and a camera motion; lower LoRAs to 0.7.

Face melts or changes identity

Pourquoi cela arrive
Face too small, too many frames, or 480p on a detailed face.

Comment le corriger
Crop closer, 49-61 frames, or 720p on the cloud; a face-detailer pass per frame is a last resort.

Camera jumps or the scene cuts

Pourquoi cela arrive
Two camera instructions, or "cut to" words in the prompt.

Comment le corriger
One camera move; put "jump cut" in the negative.

Out of memory at 720p

Pourquoi cela arrive
Frames times resolution exceed
VRAM
.

Comment le corriger
Drop to 480p or fewer frames, offload the text encoder; see the VRAM for AI video guide.

Motion LoRA does nothing

Pourquoi cela arrive
Loaded on the wrong expert or the wrong model size.

Comment le corriger
Check the model page for 5B/14B and high/low noise; load accordingly with its trigger word.

AD
PirateDiffusion
Wan 2.2 14B, LTX-2 and H3 from Telegram
PirateDiffusion runs the large video models on server GPUs: reply to any image with a workflow command and the clip comes back in the chat. Unlimited generation on a fixed price, no 24 GB card needed.

Questions

Which video model first?

Wan 2.2 I2V: forgiving, well documented, the most LoRAs. The Wan getting-started guide has the full setup.

Can I do this on an 8 GB card?

Short 480p clips with
Wan 2.1
1.3B or Wan 2.2 5B GGUF, slowly. Realistically, start on the cloud.

How do I get sound?

LTX-2 and MiniMax H3 generate synced audio; Mark White's H3 prompting guides explain dialogue and effects prompts.

How long can a clip be?

About five seconds natively. Longer pieces are chained clips; quality degrades past six to eight seconds in one pass.

Are the newest commercial models (Veo, Kling, Wan 3.0) available locally?

No; they are API-only. Everything in this guide is open-weight and runs locally or on the cloud partners.

Liens et sources

Modèles de ce guide

Écrit par
Lookout

Guides associés

Modèles vidéo
Wan 2.2 video: text-to-video and image-to-video setup, settings and LoRAs
Getting first results from Wan 2.2: which model size to pick (5B or 14B), the high-noise / low-noise expert pair, resolution and frame counts that work, the lightning LoRAs that cut renders to four steps, prompt structure for motion, and how to run it hosted when the card is too small.
Par Captain
2026-05-24
Modèles vidéo
How much VRAM do you need for AI video? Wan 2.2, HunyuanVideo 1.5 and LTX-2 by GPU size (2026)
VRAM needed to run Wan 2.2, HunyuanVideo 1.5 and LTX-2 locally, by GPU size, resolution and clip length: what 8, 12-16, 24, 32 and 48+ GB cards can make, why video costs a hundred times an image, the offloading tricks that fit 14B models on 24 GB, realistic generation times, and when renting beats buying.
Par P.I. Panda
2026-10-06
Prompting
Prompting MiniMax H3: text to video, image to video, reference video and audio
How to direct MiniMax H3 from a chat command: the WHO + WHERE + ACTION + CAMERA + SOUND recipe for text to video, animating a still with image to video, giving every reference picture a job in reference mode, and getting clean dialogue instead of mumbling. With copy-and-try prompts and the author's PDF.
Par Mark White
2026-09-20
Modèles vidéo
MiniMax H3 video prompts: the official structure
How MiniMax wants an H3 prompt written: the alignment line for image-anchored videos, the three core fields, shots and cuts, camera moves, speaker IDs and dialogue tags, the soundscape and the music. Condensed from the official prompt writing guide.
Par Captain
2026-09-14
Modèles
LTX 2.3 Crisp Enhance: sharper, more cinematic video detail
vrgamedevgirl's Crisp Enhance LoRA for LTX 2.3 increases sharpness, micro-detail and contrast in generated video. Learn how to balance it with the Soft Enhance sibling, weights and prompt ideas.
Par Captain
2026-10-03
Erreurs et solutions
GPU and VRAM guide for AI image and video: what each model needs and what to buy (or not)
How much graphics memory each model family really needs at usable speed (SD 1.5, SDXL, Flux, FLUX.2, Qwen Image, Z-Image, Wan 2.2, LTX-2, H3), what fp8 and GGUF quantisation buy you, why system RAM and disk matter too, a tier list of cards from 8 to 32 GB, Mac and AMD notes, and the point at which renting is cheaper than buying.
Par Quartermaster
2026-01-07

Autres guides

Modèles
FLUX.1 Dev: how to prompt it, best settings and creative ideas
A practical guide to FLUX.1 Dev by Black Forest Labs: natural-language prompting, guidance and step settings, LoRA stacking, text rendering and prompt ideas that play to its strengths.
Par Captain
2026-09-30
Modèles
Stable Diffusion XL 1.0: prompts, settings and what it still does best
How to get the most out of the official SDXL 1.0 base model: native resolutions, CFG and sampler settings, the refiner, negative prompts and ideas for styles where SDXL still shines.
Par Captain
2026-10-01
Modèles
Stable Diffusion 1.5: the classic model, prompted properly
Stable Diffusion 1.5 is still worth knowing: the right resolution, CFG and sampler, how to use its huge library of LoRAs and embeddings, and creative prompt ideas suited to a 512-pixel model.
Par Captain
2026-10-02
Modèles
FLUX.2 Dev: prompting the newest Black Forest Labs model
FLUX.2 Dev brings larger prompts, better text, stronger realism and multi-reference editing. This guide covers the turbo workflow, guidance and resolution settings, prompt structure and ideas that use its new strengths.
Par Captain
2026-10-03
Modèles
Qwen-Image: long prompts, perfect text and bilingual posters
Qwen-Image is the model to reach for when the words in the picture matter. Learn how to write its long descriptive prompts, render English and Chinese text accurately, set CFG and steps, and explore layout-heavy creative ideas.
Par Captain
2026-09-29
Modèles
SDXL-Lightning: four-step generation on any SDXL checkpoint
ByteDance's SDXL-Lightning LoRA turns a 30-step SDXL render into a 4-step one. Learn the step counts, the CFG you must use, the sampler, and how to combine it with your favorite checkpoints and style LoRAs.
Par Captain
2026-09-30