CUDA out of memory: how much VRAM each AI model needs and how to fit it
Topic: Errors & fixes
By Captain
Published 2026-05-18
Overview
CUDA out of memory is the most common error in local AI image and video generation, and it is not a bug: the model, the text encoders, the latent image and the attention activations must all live in the graphics card's memory at once. How much they need depends on the family, the precision of the weights, the resolution, the batch size and, for video, the number of frames.
This guide gives the numbers people actually measure, not the theoretical ones. "Comfortable" means the default workflow at the family's native resolution runs without offloading; "minimum" means it runs with quantised weights and offload tricks. Both assume a single image or clip at a time.
If a model sits above your card, the hosted route is cheaper than a new GPU:
BitVector
runs these families on cloud GPUs in a web app and Discord, and
PirateDiffusion
runs them from Telegram with the trigger words shown in this site's Run box. Many people use a local card for
SD 1.5
and
SDXL
and the hosted services for
Flux
,
Qwen
and video.
Reference
Name | Type | What it is |
|---|---|---|
4 GB min / 6 GB comfortable | 512x512 native. 2 GB file (fp16). Runs on almost anything; use --lowvram on 4 GB. | |
6 GB min / 8-12 GB comfortable | ||
8 GB min (GGUF Q5 / NF4) / 16 GB comfortable (fp8) / 24 GB full | 1024x1024. 23 GB fp16 file; 11 GB fp8; 6.8 GB NF4. T5 encoder adds 5 GB fp16 or 2.5 GB fp8. | |
16 GB min (GGUF) / 24 GB comfortable (fp8) / 32 GB+ full | Up to 4 MP. Larger text encoder; fp8 is the practical choice on consumer cards. | |
Qwen-Image / Qwen-Image-Edit | 12 GB min (GGUF Q4) / 24 GB comfortable (fp8) | |
8 GB min / 16 GB comfortable | 6B parameters, 8 steps, 1024 native. The friendliest modern model for 8-12 GB cards. | |
Wan 2.2 5B (TI2V) | 8 GB min / 12 GB comfortable | 720p, 5 seconds. The 5B hybrid is built for consumer cards. |
Wan 2.2 14B (T2V / I2V) | 12 GB min (GGUF Q4 + offload) / 24 GB comfortable (fp8) | 480p-720p, 81 frames. Two expert models (high / low noise) load in turn. Lightning LoRAs make 4 steps usable. |
12 GB min (GGUF + offload) / 24 GB comfortable | 13B, 544x960, 65-129 frames. Frame count is the lever. | |
12 GB min / 24 GB comfortable | Audio and video together; 4K needs 24 GB+. Fast per frame. |
AD

No install required - run AI models on the cloud
BitVector Prism is the easiest way to start: pick a model, type a prompt and generate in a clean web app, with nothing to configure. BitVector is also available on Discord and on the web (SpyGlass).
Step by step
- Find the family and size of the model on its page here (family at the top, file sizes under Files) and compare with the table above and your card.
- Pick the precision the card can hold: fp16 only if the file plus 4 GB fits; otherwise fp8 (same quality for images), then GGUF Q8, Q6, Q5, Q4 in that order. ComfyUI-GGUF provides the Unet and CLIP loaders.
- Quantise the text encoder too: the T5-XXL fp8 file is half the size of fp16 with no visible change for most prompts.
- Keep resolution at the native size and upscale afterwards: 512 for SD 1.5, 1024 for SDXL, Flux, Qwen andZ-Image. Doubling the side quadruples the activation memory.
- Batch size 1. Generate more images in sequence instead of in parallel.
- Decode in tiles: VAE Decode (Tiled) in ComfyUI, Tiled VAE in A1111/Forge. The decode step alone can fail on a sampled latent that fit.
- Video: fewer frames first (33 or 49), then shorter sides (480p), then more steps. Lightning / distillation LoRAs listed on the Wan andHunyuanfamily pages cut steps from 30 to 4 to 8.
- Still short: run it hosted. The Run box on each model page opens the model on PirateDiffusion (Telegram) or BitVector (web, Discord) with nothing to install.
Examples
ComfyUI Flux dev on 8 GB
Unet Loader (GGUF): flux1-dev-Q5_K_M.gguf
DualCLIPLoader (GGUF): t5-v1_1-xxl-encoder-Q5_K_M.gguf + clip_l.safetensors
VAE: ae.safetensors | 1024x1024 | 20 steps | --lowvram
Wan 2.2 14B I2V on 12 GB
High noise: wan2.2_i2v_high_noise_14B_Q4_K_M.gguf | Low noise: wan2.2_i2v_low_noise_14B_Q4_K_M.gguf
umt5_xxl_fp8_e4m3fn_scaled.safetensors | 480x832 | 49 frames | lightning LoRA, 4 steps
Forge SDXL on 6 GB
UI: xl | GPU Weights: 4000 | Diffusion in Low Bits: Automatic | 1024x1024 | batch 1
Tips
- System RAM matters almost as much as VRAM once you offload: 32 GB is the comfortable minimum for Flux and Wan, 64 GB forFLUX.2and Qwen fp16.
- Close the browser's hardware acceleration tabs and anything using the GPU (game launchers, video calls): 1 to 2 GB of VRAM disappears into them.
- Windows reserves VRAM for the desktop; a second cheap card for the display frees the whole main card for generation.
- GGUF Q8 is visually identical to fp16 for images; Q5_K_M is the sweet spot for 8 GB cards; below Q4 details degrade.
- fp8 on RTX 40-series and newer is fast; on RTX 30 and older it works but weights are upcast on the fly, so GGUF can be quicker there.
- Keep a 512x512 or 1024x1024 test prompt handy and raise one variable at a time (resolution, then LoRAs, then ControlNet) to find the ceiling of your card.
- The heat ranking on this site shows which quantised and distilled releases the community actually uses; the lightning LoRAs for Wan and Qwen are usually near the top.
Troubleshooting
Out of memory before the first step
Why it happens
The weights alone exceed VRAM.
How to fix it
A smaller precision (fp8, GGUF Q5) or offload. Check file size versus VRAM minus 2 GB.
Out of memory in the sampler at high resolution
Why it happens
Attention activations scale with pixel count.
How to fix it
Out of memory only at VAE decode
Why it happens
Decoding is the memory peak for large images and for video.
How to fix it
Tiled VAE decode; fp16/bf16 VAE flag; for video decode in temporal chunks.
Fits once, fails the second time
Why it happens
Fragmentation and cached models from the previous run.
How to fix it
Unload models (ComfyUI Manager has a button), or set PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True before start.
Slow but no error after adding --lowvram
Why it happens
Weights stream from system RAM every step.
How to fix it
Expected. Reduce the model (quantise) so more stays on the GPU, or run hosted for the big families.
System freezes or reboots
Why it happens
System RAM exhausted and the page file is too small, or a power supply at its limit.
How to fix it
System-managed page file on an SSD; a power limit of 80 to 90 percent via the driver tool costs little speed.
AD

No install required - run AI models on the cloud
PirateDiffusion is Telegram only and built for pros: thousands of models, LoRAs and workflows driven by chat commands, with unlimited generation on a fixed price plan.
Questions
Is 8 GB enough in 2026?
For SD 1.5, SDXL,
Z-Image Turbo
and
Flux dev
in GGUF, yes. For FLUX.2,
Qwen-Image
and the 14B video models it is the minimum with heavy offload, and generation is slow; most people run those hosted.
Does fp8 lower quality?
For image models, not visibly. GGUF Q8 is equivalent; Q5 shows tiny differences in fine text and textures; Q4 loses some detail.
Why does a 6.5 GB SDXL file need 8 GB?
Text encoders, VAE, the latent and attention activations come on top of the weights, and Windows keeps some VRAM for the display.
What about Apple Silicon or AMD?
Unified memory on
Apple Silicon
makes 32 GB machines run Flux fp8, slowly; AMD runs through ROCm on Linux or ZLUDA / DirectML on Windows, with less headroom. Hosted services sidestep both.
Which is cheaper, a 24 GB card or a hosted plan?
A hosted plan for a year costs less than a used 3090 and covers every family at once; a card pays off if you generate for hours every day. Many people do both.
Links and sources
- ComfyUI-GGUF loaders
- PyTorch memory management notes
- Run the big models hosted on PirateDiffusion
- BitVector cloud GPUs
Models in this guide
Stable Diffusion 1.5 Models
Stable Diffusion XL 1.0 Models
Flux Models
Flux 2 / Klein Models
Wan Models
Qwen Models
Hunyuan Models
Zimage Models
Written by
Captain
Related guides
Errors & fixes
ComfyUI errors and how to fix them: red nodes, missing models, CUDA out of memory
The ComfyUI error messages people hit first, what each one means and the fix that works: missing custom nodes (red boxes), "Prompt outputs failed validation", CUDA out of memory, wrong model type in a loader, mat1 and mat2 shape errors, header deserialization, torch and xformers mismatches.
By Captain
2026-05-12
Errors & fixes
Stable Diffusion WebUI Forge errors and fixes (Flux, SDXL, low VRAM)
Forge-specific problems and their fixes: Flux not loading or producing noise, the UI type switch (sd / xl / flux), "GPU Weights" slider and out-of-memory errors, missing text encoders and VAE for Flux, ControlNet and extension breakage, and the differences from Automatic1111 that trip people up.
By Captain
2026-05-16
Errors & fixes
Automatic1111 (Stable Diffusion WebUI) errors and fixes
Fixes for the Stable Diffusion WebUI errors that stop most installs: "Torch is not able to use GPU", NansException, CUDA out of memory, "Couldn't install torch", xformers, "Stable diffusion model failed to load", broken extensions after an update, and the SDXL-specific problems.
By Captain
2026-05-14
Video models
Wan 2.2 video: text-to-video and image-to-video setup, settings and LoRAs
Getting first results from Wan 2.2: which model size to pick (5B or 14B), the high-noise / low-noise expert pair, resolution and frame counts that work, the lightning LoRAs that cut renders to four steps, prompt structure for motion, and how to run it hosted when the card is too small.
By Captain
2026-05-24
Models
Flux vs SDXL vs SD 1.5 vs Qwen vs Z-Image: which model family to use in 2026
A practical comparison of the open image model families: prompt following, realism, anime and illustration, text rendering, speed, VRAM, LoRA ecosystem and licence, with a recommendation for each kind of work and hardware.
By Captain
2026-05-22
More guides
Cloud services
Flat fee vs tokens: why unlimited plans like Graydient.ai are the best value for AI creators in 2026
A ranked cost comparison, with prices checked on 8 October 2026, of flat-fee unlimited plans like Graydient.ai against token and credit pricing from Midjourney, Leonardo, fal.ai, Replicate, Runway and Kling: what 1,000 images and 1,000 videos a month really cost on each, and why one fixed fee for unlimited images, video, audio, Grok LLM chat and web apps is the best value for creators who iterate.
By Captain
2026-10-08
Models
FLUX.1 Dev: how to prompt it, best settings and creative ideas
A practical guide to FLUX.1 Dev by Black Forest Labs: natural-language prompting, guidance and step settings, LoRA stacking, text rendering and prompt ideas that play to its strengths.
By Captain
2026-09-30
Models
Stable Diffusion XL 1.0: prompts, settings and what it still does best
How to get the most out of the official SDXL 1.0 base model: native resolutions, CFG and sampler settings, the refiner, negative prompts and ideas for styles where SDXL still shines.
By Captain
2026-10-01
Models
Stable Diffusion 1.5: the classic model, prompted properly
Stable Diffusion 1.5 is still worth knowing: the right resolution, CFG and sampler, how to use its huge library of LoRAs and embeddings, and creative prompt ideas suited to a 512-pixel model.
By Captain
2026-10-02
Models
FLUX.2 Dev: prompting the newest Black Forest Labs model
FLUX.2 Dev brings larger prompts, better text, stronger realism and multi-reference editing. This guide covers the turbo workflow, guidance and resolution settings, prompt structure and ideas that use its new strengths.
By Captain
2026-10-03
Models
Qwen-Image: long prompts, perfect text and bilingual posters
Qwen-Image is the model to reach for when the words in the picture matter. Learn how to write its long descriptive prompts, render English and Chinese text accurately, set CFG and steps, and explore layout-heavy creative ideas.
By Captain
2026-09-29

Model
Trends
.ai
© ModelTrends.ai
|
© 2026
|
All Rights Reserved