VRAM requirements for every popular AI image and video model (2026): full precision, FP8 and GGUF
Chủ đề: Lỗi và cách khắc phục
Bởi P.I. Panda
Đăng ngày 2026-10-06
How much VRAM you need for SD 1.5, SDXL, SD 3.5, Z-Image Turbo, FLUX.1, FLUX.2 Klein and Dev, Qwen-Image, Wan 2.2, HunyuanVideo 1.5 and LTX-2, at full precision, FP8 and 4-bit GGUF, with the comfortable card and licence for each, what every GPU size from 8 to 48 GB can run, and the order of fixes when a model does not fit.
Trang này chưa được dịch nên đang hiển thị bằng tiếng Anh.
Tổng quan
Short answer: 8 GB of
VRAM
runs
SDXL
and the small fast models such as
Z-Image Turbo
and
FLUX.2 Klein 4B
. 12 GB runs
FLUX.1
and
Qwen-Image
once they are quantised. 24 GB runs almost every image model plus
Wan 2.2
video in FP8. The newest large models,
FLUX.2 Dev
and
LTX-2
at full quality, want 32 GB or more.
The table below lists every popular open model with the memory it needs at full precision, at FP8 and as a 4-bit GGUF file, generating at about 1024x1024. Different testers count the text encoder differently, so treat every number as a range rather than a hard limit. Three things the table hides: SDXL is still the cheapest model to run well (the base fits in 8 GB, but
LoRAs
,
ControlNet
and
upscaling
add memory, which is why 12 GB is the comfortable number); FLUX.2 Klein 4B is the best fit for small cards right now (about 13 GB at full precision, roughly half in FP8, Apache 2.0); and FLUX.2 Dev is not a home model at full precision (one
ComfyUI
user ran FP8 on a 12 GB RTX 3060, using about 70 GB of system RAM to get there).
Precision changes everything. BF16/FP16 is full quality and full size, about 2 GB of VRAM per billion parameters. FP8 halves that, about 1 GB per billion, and most people cannot tell FP8 output from full precision. GGUF files (Q8, Q5, Q4 and lower) shrink models further: Q8 is close to full quality, Q4 trades some detail for a much smaller footprint, and below Q4 quality drops quickly. The model is also not the only thing in memory: text encoders (T5 for FLUX.1, a language model for Qwen-Image and
FLUX.2
), the VAE, LoRAs, ControlNet and the image itself all take space, and ComfyUI moves some of these to system RAM automatically, which is why two people with the same card report different numbers. Figures checked 6 October 2026 against model cards and community tests; updated monthly.
Tham khảo
Model
Full precision
FP8
4-bit / GGUF
Comfortable card
Stable Diffusion 1.5
0.9B
CreativeML OpenRAIL-M. Runs on almost anything.
Full precision
4 GB
FP8
n/a
4-bit / GGUF
n/a
Comfortable card
6-8 GB
SDXL and fine-tunes: Pony, Illustrious, NoobAI
3.5B
OpenRAIL++ (fine-tunes vary). 12 GB leaves room for LoRAs, ControlNet and upscaling in one workflow.
Full precision
6-8 GB
FP8
n/a
4-bit / GGUF
n/a
Comfortable card
12 GB
SD 3.5 Large
8.1B
Stability Community License.
Full precision
16-24 GB
FP8
lower
than full precision
4-bit / GGUF
lower
than FP8
Comfortable card
24 GB
Z-Image Turbo
6B
Apache 2.0. 8 steps; the fast model for small cards.
Full precision
~16 GB
FP8
~8 GB
4-bit / GGUF
~6 GB
Comfortable card
8-12 GB
FLUX.2 Klein 4B
4B
Apache 2.0, commercial use allowed. Best fit for 8 GB cards in 2026.
Full precision
~13 GB
FP8
~8 GB
4-bit / GGUF
~4-5 GB
Comfortable card
8-12 GB
FLUX.2 Klein 9B
9B
FLUX Non-Commercial License (outputs usable commercially).
Full precision
~29 GB
FP8
~17 GB
4-bit / GGUF
~8-9 GB
Comfortable card
16-24 GB
FLUX.1 Schnell
12B
Apache 2.0. 4 steps.
Full precision
~24 GB
FP8
~12 GB
4-bit / GGUF
~6-8 GB
Comfortable card
12-16 GB
FLUX.1 Dev
12B
Non-commercial licence. The fp8 T5 encoder saves about 4 GB on its own.
Full precision
~24 GB
more with the text encoder loaded
FP8
~12-18 GB
4-bit / GGUF
~7 GB
Comfortable card
16-24 GB
Qwen-Image
20B
Apache 2.0. Strong text rendering; needs its Qwen2.5-VL encoder in FP8 on anything under 32 GB.
Full precision
~41 GB
FP8
~20 GB
4-bit / GGUF
~12-13 GB
Comfortable card
24 GB
FLUX.2 Dev
32B
Non-commercial licence. Expect slow generations without a 32 GB card.
Full precision
~90 GB
with its text encoder
FP8
32 GB
fits a 32 GB card
4-bit / GGUF
12-16 GB
heavy offloading
Comfortable card
32 GB+
HunyuanImage 3.0
80B MoE
Tencent Community License. Datacenter only; use it through a hosted service.
Full precision
multi-GPU
at every precision
FP8
n/a
4-bit / GGUF
n/a
Comfortable card
datacenter
Wan 2.1 T2V 1.3B
1.3B
Video 480p
Apache 2.0.
Full precision
8.19 GB
official
FP8
8 GB
in practice
4-bit / GGUF
n/a
Comfortable card
12 GB
Wan 2.2 TI2V 5B
5B
Video 720p
Apache 2.0.
Full precision
consumer GPU
official, tested on an RTX 4090
FP8
n/a
4-bit / GGUF
8 GB
GGUF Q8
Comfortable card
24 GB
Wan 2.2 A14B T2V/I2V
2 x 14B
Video 720p
Apache 2.0. Two 14B experts, one in VRAM at a time, which is why 24 GB cards handle it. Wan 2.5-3.0 are API only.
Full precision
80 GB
official, single GPU
FP8
~14 GB
per expert
4-bit / GGUF
12-16 GB
GGUF
Comfortable card
24-32 GB
HunyuanVideo 1.5
8.3B
Video 720p, 1080p with its upscaler
Tencent Community License.
Full precision
14 GB
official, with offloading
FP8
14 GB
in practice
4-bit / GGUF
n/a
Comfortable card
24 GB+
80 GB for best quality
LTX-2 / LTX-2.3
19B / 22B
Video up to 4K
LTX-2 Community License. Synced audio.
Full precision
32 GB+
official
FP8
n/a
4-bit / GGUF
24 GB
FP4
Comfortable card
32 GB
AD

Every row of this table, no VRAM required
BitVector Prism runs SDXL, Z-Image, FLUX.1, FLUX.2 and Qwen-Image at full precision on its own GPUs. Pick a model from this site, type a prompt, generate in the browser. Nothing to download.
Từng bước
- Find your VRAM (Task Manager > Performance > GPU, or nvidia-smi) and match it to a tier below; then pick models whose "comfortable card" column is at or below it.
- 12 GB (RTX 3060 12GB, 4070, 5070): everything above, faster, with room for ControlNet and upscaling;FLUX.1 Devand Schnell in FP8 or GGUF; Qwen-Image in GGUF Q4;FLUX.2 Klein 9Bin GGUF; Wan 2.2 A14B in GGUF Q4 with offloading at low resolution and short clips.
- 16 GB (RTX 4060 Ti 16GB, 4070 Ti Super, 5070 Ti, 5080): FLUX.1 Dev in FP8 with headroom; Z-Image Turbo and Klein 4B at full precision; Qwen-Image in GGUF Q4-Q5;HunyuanVideo1.5 with offloading; Wan 2.2 A14B in GGUF Q5 at 720p (community reports).
- 24 GB (RTX 3090, 4090): almost every image model except FLUX.2 Dev at full precision; Qwen-Image in FP8 or Q8; Wan 2.2 A14B in FP8 and Wan 2.2 5B at full precision; LTX-2 in FP4; most LoRA training for SDXL and Flux.
- 32 GB (RTX 5090): FLUX.2 Dev in FP8; LTX-2 andLTX-2.3in FP8; Wan 2.2 A14B at 720p without runningout of memorypartway through a clip.
- 48 GB and up (RTX 6000 Ada, A100, H100): full-precision Qwen-Image, LTX-2 and FLUX.2 Dev with room to spare; long or high-resolution video, batch generation and serious training.
- Download the precision for your tier from the model page on this site (it lists each variant with its file size), load text encoders in FP8, and keep 32 GB of system RAM (64 GB for the largest models) for offloading.
Ví dụ
FLUX.1 Dev on a 12 GB card (ComfyUI)
Load Diffusion Model: flux1-dev-fp8-e4m3fn (11.9 GB) or UnetLoaderGGUF flux1-dev-Q6_K (9.8 GB) | DualCLIPLoader: clip_l + t5xxl_fp8_e4m3fn_scaled | Load VAE: ae
1024x1024, 20 steps, euler, simple, guidance 3.5 | VAE Decode (Tiled) if the last step fails
Qwen-Image on 16 GB
UnetLoaderGGUF: qwen_image_Q4_K_M.gguf (~12.5 GB) | CLIPLoader: qwen_2.5_vl_7b_fp8_scaled, device cpu | Load VAE: qwen_image_vae
1328x1328 native; start at 1024x1024, 20 steps, cfg 2.5
The same models with no card at all
/render <flux-dev> a studio photo of a vintage camera on a walnut desk, soft window light /size:1024x1024 (PirateDiffusion, Telegram)
Mẹo
- When you see CUDA out of memory, do these in order: switch to an FP8 or GGUF version of the same model (the biggest single saving); lower the resolution or clip length (memory grows with pixel count, 1536 to 1024 px cuts a lot); use tiled VAE decoding; load the text encoder in FP8 or keep it on the CPU; close other GPU apps (browsers and games hold VRAM too).
- Offloading works but you pay in speed and system RAM: a model can run on a card "too small" for it, several times slower, and you may need 32-64 GB or more of system RAM.
- Licences in one line: Z-Image Turbo, FLUX.2 Klein 4B, FLUX.1 Schnell, Qwen-Image andWan2.1/2.2 are Apache 2.0 (commercial use allowed); FLUX.1 Dev, FLUX.2 Dev and Klein 9B use Black Forest Labs' non-commercial licence, which forbids using the model in a commercial product or service but lets you use the generated images commercially. Licences change; the model page links the current one.
- Buying a 32 GB card to run FLUX.2 Dev or LTX-2 for a weekend project rarely pays off; if the model you want sits in the 24 GB+ rows, renting by the hour or a flat-rate cloud plan is usually cheaper and faster than offloading on a small card.
- Per-billion-parameter rule: 2 GB at BF16, 1 GB at FP8, about 0.6 GB at Q4, plus 2-5 GB for the encoder, VAE and latent.
- For video, resolution and clip length matter as much as the model; the companion guide onVRAM for AI videobreaks that down.
Khắc phục sự cố
CUDA out of memory on a card the table says is fine
Vì sao xảy ra
Full-precision text encoder, a LoRA stack, ControlNet or a high resolution on top of the model.
Cách sửa
FP8 encoder, one LoRA, native resolution, tiled VAE; the CUDA out of memory guide has the long version.
Generation is painfully slow although nothing crashes
Vì sao xảy ra
Automatic offloading to system RAM.
Cách sửa
Drop to the next precision (FP8 to Q6, Q6 to Q4) until Shared GPU memory stays near zero.
FLUX.2 Dev "works" on 12 GB but needs 70 GB of RAM
Vì sao xảy ra
FP8 weights swapped from system memory.
Cách sửa
Use Klein 4B/9B locally and FLUX.2 Dev on the cloud or a 32 GB card.
Cannot tell FP8 from full precision, so why does Q4 look soft?
Vì sao xảy ra
Q4 discards more detail than FP8; below Q4 quality drops fast.
Cách sửa
Use Q6 or Q8 when they fit; Q4 for drafts.
Licence confusion on FLUX
Vì sao xảy ra
Schnell and Klein 4B are Apache 2.0; Dev, FLUX.2 Dev and Klein 9B are non-commercial.
Cách sửa
Check the licence column and the model card before using a model in a product.
AD

The 24 GB+ rows, from your phone
PirateDiffusion hosts FLUX.2 Dev, Qwen-Image, Wan 2.2 14B and LTX-2 on server hardware and runs them from Telegram commands. Unlimited generation on a fixed monthly price; no card, no offloading, no 70 GB of RAM.
Câu hỏi
Can I run Flux on 8 GB of VRAM?
Yes, with a quantised file. FLUX.1 in GGUF Q4 needs about 7 GB, and FLUX.2 Klein 4B in FP8 about 8 GB. Expect slower generations than on a 12 or 16 GB card. Full-precision FLUX.1 needs about 24 GB.
Is 12 GB enough for SDXL?
Yes. SDXL itself fits in 8 GB; 12 GB gives you room for LoRAs, ControlNet and upscaling in the same workflow.
How much VRAM does Wan 2.2 need?
It depends on the version. The 5B model runs on 8 GB as a GGUF file and comfortably on 24 GB. The A14B model officially needs 80 GB without offloading, but its FP8 version runs well on 24 GB and GGUF versions run on 12-16 GB with offloading.
Does system RAM matter?
Yes, once you start offloading. 32 GB is a sensible minimum for Flux and Wan, and the largest models (FLUX.2 Dev on a 12 GB card, for example) can use 64 GB or more.
Which models can I use commercially?
Z-Image Turbo, FLUX.2 Klein 4B, FLUX.1 Schnell, Qwen-Image and Wan 2.1/2.2 (Apache 2.0). FLUX.1 Dev, FLUX.2 Dev and Klein 9B are non-commercial for the model itself, though the images you generate may be used commercially. Always check the model card.
Is there a newer Wan I can run at home?
No. Wan 2.2 is the newest Wan release with downloadable weights; Wan 2.5, 2.6, 2.7 and 3.0 are API-only.
Liên kết và nguồn
- Wan 2.2 A14B model card
- FLUX.2 Klein 9B licence
- ComfyUI GGUF loaders
- Flux family page
- FLUX.2 family page
- Qwen Image family page
- Z-Image family page
- Wan family page
- PirateDiffusion
- BitVector web app
Mô hình trong hướng dẫn này
Mô hình Stable Diffusion 1.5
Mô hình Stable Diffusion XL 1.0
Mô hình Zimage
Mô hình Flux
Mô hình Flux 2 / Klein
Mô hình Qwen
Mô hình Wan
Mô hình Hunyuan
Mô hình Ltx2
Người viết
P.I. Panda
Hướng dẫn liên quan

Mô hình video
How much VRAM do you need for AI video? Wan 2.2, HunyuanVideo 1.5 and LTX-2 by GPU size (2026)
VRAM needed to run Wan 2.2, HunyuanVideo 1.5 and LTX-2 locally, by GPU size, resolution and clip length: what 8, 12-16, 24, 32 and 48+ GB cards can make, why video costs a hundred times an image, the offloading tricks that fit 14B models on 24 GB, realistic generation times, and when renting beats buying.
Bởi P.I. Panda
2026-10-06
Lỗi và cách khắc phục
CUDA out of memory: how much VRAM each AI model needs and how to fit it
A VRAM table for SD 1.5, SDXL, Flux, FLUX.2, Qwen-Image, Z-Image, Wan 2.2 and HunyuanVideo, and the techniques that make a model fit: fp8 and GGUF quantisation, CPU offload, tiled VAE, resolution and frame limits, and when to stop fighting and run it hosted.
Bởi Captain
2026-05-18
Lỗi và cách khắc phục
GPU and VRAM guide for AI image and video: what each model needs and what to buy (or not)
How much graphics memory each model family really needs at usable speed (SD 1.5, SDXL, Flux, FLUX.2, Qwen Image, Z-Image, Wan 2.2, LTX-2, H3), what fp8 and GGUF quantisation buy you, why system RAM and disk matter too, a tier list of cards from 8 to 32 GB, Mac and AMD notes, and the point at which renting is cheaper than buying.
Bởi Quartermaster
2026-01-07
Mô hình
Flux vs SDXL vs SD 1.5 vs Qwen vs Z-Image: which model family to use in 2026
A practical comparison of the open image model families: prompt following, realism, anime and illustration, text rendering, speed, VRAM, LoRA ecosystem and licence, with a recommendation for each kind of work and hardware.
Bởi Captain
2026-05-22
Lỗi và cách khắc phục
ComfyUI errors and how to fix them: red nodes, missing models, CUDA out of memory
The ComfyUI error messages people hit first, what each one means and the fix that works: missing custom nodes (red boxes), "Prompt outputs failed validation", CUDA out of memory, wrong model type in a loader, mat1 and mat2 shape errors, header deserialization, torch and xformers mismatches.
Bởi Captain
2026-05-12
Dịch vụ đám mây
PirateDiffusion: run any model from Telegram, no GPU needed
The commands that matter in the PirateDiffusion Telegram bot: /render with model trigger words, (( )) and [[ ]] weighting, #recipes, /adetailer, /highdef and /facelift upscaling, /remix, /inpaint, ComfyUI workflows with /wf /run:, and the new // skills that pick the model for you.
Bởi Captain
2026-10-03
Ứng dụng
BitVector: Vector for Discord, Spyglass on the web and Prism chat
How to run hosted models and ComfyUI workflows from Discord with the Vector bot, from any browser with Spyglass, and by chatting with Prism: account linking, /ai render and /ai workflow, initial media for image-to-image and image-to-video, skills, and the fixes for the errors people hit first.
Bởi Captain
2026-10-03
Hướng dẫn khác
Mô hình
FLUX.1 Dev: how to prompt it, best settings and creative ideas
A practical guide to FLUX.1 Dev by Black Forest Labs: natural-language prompting, guidance and step settings, LoRA stacking, text rendering and prompt ideas that play to its strengths.
Bởi Captain
2026-09-30
Mô hình
Stable Diffusion XL 1.0: prompts, settings and what it still does best
How to get the most out of the official SDXL 1.0 base model: native resolutions, CFG and sampler settings, the refiner, negative prompts and ideas for styles where SDXL still shines.
Bởi Captain
2026-10-01
Mô hình
Stable Diffusion 1.5: the classic model, prompted properly
Stable Diffusion 1.5 is still worth knowing: the right resolution, CFG and sampler, how to use its huge library of LoRAs and embeddings, and creative prompt ideas suited to a 512-pixel model.
Bởi Captain
2026-10-02
Mô hình
FLUX.2 Dev: prompting the newest Black Forest Labs model
FLUX.2 Dev brings larger prompts, better text, stronger realism and multi-reference editing. This guide covers the turbo workflow, guidance and resolution settings, prompt structure and ideas that use its new strengths.
Bởi Captain
2026-10-03
Mô hình
Qwen-Image: long prompts, perfect text and bilingual posters
Qwen-Image is the model to reach for when the words in the picture matter. Learn how to write its long descriptive prompts, render English and Chinese text accurately, set CFG and steps, and explore layout-heavy creative ideas.
Bởi Captain
2026-09-29
Mô hình
SDXL-Lightning: four-step generation on any SDXL checkpoint
ByteDance's SDXL-Lightning LoRA turns a 30-step SDXL render into a 4-step one. Learn the step counts, the CFG you must use, the sampler, and how to combine it with your favorite checkpoints and style LoRAs.
Bởi Captain
2026-09-30

Model
Trends
.ai
© ModelTrends.ai
|
Làm tại Nhật Bản
|
© 2026