Model
Trends
.ai
Model
Trends
.ai
the best open source ai models

GPU and VRAM guide for AI image and video: what each model needs and what to buy (or not)

Topic: Errors & fixes
By Quartermaster
Published 2026-01-07
How much graphics memory each model family really needs at usable speed (SD 1.5, SDXL, Flux, FLUX.2, Qwen Image, Z-Image, Wan 2.2, LTX-2, H3), what fp8 and GGUF quantisation buy you, why system RAM and disk matter too, a tier list of cards from 8 to 32 GB, Mac and AMD notes, and the point at which renting is cheaper than buying.

Overview

The single number that decides what you can run locally is VRAM, the memory on the graphics card. A model must fit in it, together with its text encoder, the VAE and the working latent. When it does not fit, UIs spill to system RAM and speed drops ten to fifty times, which is where most "it takes 20 minutes per image" complaints come from. Speed itself (how many images per minute once it fits) depends on the GPU generation, and matters less than fitting.
Requirements grew fast.
SD 1.5
was comfortable in 4 GB;
SDXL
wants 8;
Flux dev
in its original form wants 24 and in fp8 or GGUF form 12 (usable in 8);
FLUX.2
and
Qwen Image
are 20B-class models that need 16 GB with quantisation;
Wan 2.2
14B video and
LTX-2
want 24 GB for anything beyond short low-resolution clips, and the audio-video models (H3) are cloud-only for most people. Quantisation (fp8, NF4, GGUF Q4-Q8) is what keeps new models reachable on mid-range cards: smaller weights, slightly lower quality, same prompts.
This is also a buying decision. A used 24 GB RTX 3090 is still the value pick for images; a 5090 at 32 GB is the local ceiling; anything with 8 GB is an SDXL machine. Against that stands renting: a subscription to
PirateDiffusion
or
BitVector
costs less per month than the electricity of a 3090 and includes every model on this site, which is why many people run SDXL locally and everything heavier on the cloud.

Reference

Name
Type
What it is
SD 1.5
4 GB min / 6 GB comfortable
512x768, with
LoRAs
and
ControlNet
. Runs on almost anything, including older laptops.
SDXL (incl. Pony, Illustrious)
8 GB min / 12 GB comfortable
1024x1024 with a LoRA or two; 8 GB needs fp16 VAE fix and no ControlNet stacking.
Flux dev / Krea / SRPO
12 GB fp8 / 8 GB GGUF Q4 / 24 GB fp16
T5 encoder in fp8 saves 4 GB on its own. Q4 is visibly softer; Q6 or Q8 near fp8 quality.
FLUX.2 dev / Qwen Image / Qwen Image Edit
16 GB (fp8 or Q4) / 24 GB comfortable
20B-class image models.
Klein 4B
and 9B variants of FLUX.2 run in 8 to 12 GB.
Z-Image Turbo / Chroma
8-12 GB
6B-9B models; Turbo at 8 steps is fast even on 8 GB with offloading.
Wan 2.2 5B (TI2V)
12 GB
720p short clips with offloading; the entry point for local video.
Wan 2.2 14B, LTX-2, Hunyuan Video
16 GB Q4 (slow) / 24 GB / 32 GB best
Two 14B experts for Wan 2.2; block swapping and GGUF make 16 GB possible at low resolution and minutes per clip.
System RAM and disk
32 GB RAM / 64 GB for video; NVMe
Offloading and model switching live in RAM; models load from disk every switch. 1 TB fills in a month.
AD
No GPU? Skip this whole guide
BitVector Prism runs SDXL, Flux, FLUX.2, Qwen and Z-Image on its own hardware; you need a browser. Try the models from this site at full precision before deciding whether to buy a card at all.

Step by step

  1. Find your VRAM: Windows Task Manager > Performance > GPU > Dedicated GPU memory; Linux nvidia-smi; Mac uses unified memory (the whole RAM, shared).
  2. Match it to the table: 8 GB means SDXL natively and Flux/Z-Image via GGUF; 12 GB adds
    Flux
    fp8 and small video; 16 GB adds FLUX.2/Qwen quantised; 24 GB+ runs everything except the largest video at full quality.
  3. Download the quantised build the model page lists for your tier (fp8 for 12-16 GB, GGUF Q4-Q6 for 8-12 GB) rather than the fp16 original.
  4. Set the UI to offload:
    ComfyUI
    --lowvram or --novram flags (or automatic),
    Forge
    's GPU weights slider, --medvram in A1111.
  5. Use fp8 text encoders (t5xxl_fp8,
    qwen
    2.5 vl fp8) and tiled VAE decode; these are the two biggest savings after the model itself.
  6. Close the browser GPU acceleration, games and a second monitor's heavy apps; 1 to 2 GB often goes there.
  7. If a model still does not fit at usable speed, run that family on the cloud and keep the local card for what it does well.
  8. Buying: 24 GB used (3090) or new (4090/5090 class) for serious local video; 16 GB (4070 Ti Super / 5070 Ti class) for images including FLUX.2 and Qwen; 12 GB as the floor for Flux fp8; avoid 8 GB cards for anything bought new in 2026.

Examples

ComfyUI flags for a 8 GB card

python main.py --lowvram --use-split-cross-attention
Flux: flux1-dev-Q4_K_S.gguf + t5xxl_fp8_e4m3fn + clip_l + ae.safetensors, 1024x1024, 20 steps

Reading a model page for memory

Variant list: fp16 23.8 GB | fp8 11.9 GB | Q8 12.7 GB | Q6_K 9.8 GB | Q4_K_S 6.8 GB
Rule of thumb: pick the largest variant that leaves 3 GB free after the text encoder (4.9 GB fp8 T5 for Flux)

Tips

  • VRAM beats speed. A 24 GB older card runs more models than a faster 12 GB card.
  • NVIDIA is the default: CUDA support is universal. AMD works through ROCm on Linux and is improving on Windows, with fewer custom nodes supported; Intel Arc runs via IPEX with similar caveats.
  • Apple Silicon
    shares RAM with the GPU: a 32 GB Mac runs Flux fp8 slowly but without out-of-memory; a 16 GB Mac is an SDXL machine.
  • Quantised models differ: fp8 (needs a 40-series or newer for native speed), NF4 (Forge, fast on small cards), GGUF (ComfyUI, finest size control). Pick what your UI supports.
  • Video memory use scales with frames x resolution; halve the frame count before lowering the model precision.
  • The model pages on this site show the file size per variant; as a rule of thumb you need the file size plus 2 to 4 GB in VRAM.

Troubleshooting

CUDA out of memory

Why it happens
Model + encoder + latent exceed VRAM.

How to fix it
Quantised build, fp8 encoder, tiled VAE, smaller size or batch; our
CUDA out of memory
guide has the full checklist.

Generation suddenly takes minutes

Why it happens
Shared GPU memory (system RAM) is being used because the model does not fit.

How to fix it
Check Task Manager "Shared GPU memory"; reduce footprint until it stays at zero.

fp8 model is slow on my 30-series card

Why it happens
Native fp8 compute needs Ada (40-series) or newer; older cards convert on the fly.

How to fix it
Use GGUF Q8 or NF4 instead on 20- and 30-series cards.

Video model works at 480p but not 720p

Why it happens
Latent size grows with the square of the resolution times frames.

How to fix it
Fewer frames, lower resolution and upscale after, or run 720p on the cloud.

Mac runs out of memory at 64 GB

Why it happens
macOS caps GPU-allocatable memory below the total, and fp16 is not always available on MPS.

How to fix it
Raise the wired limit (sysctl iogpu.wired_limit_mb), use fp8/GGUF weights, smaller batches.

AD
PirateDiffusion
Video and 20B models without a 24 GB card
PirateDiffusion runs Wan 2.2, LTX-2, FLUX.2 and Qwen Image on server GPUs and sends the result to your Telegram. Fixed monthly price, unlimited generation, less than the power bill of a 3090.

Questions

Is 8 GB enough in 2026?

For SDXL,
Z-Image Turbo
and GGUF Flux, yes. For FLUX.2, Qwen Image and video, you will be waiting or on the cloud.

Laptop GPU or desktop?

Laptop cards have the same VRAM numbers but throttle; a 16 GB laptop card is fine for images, poor for video. Desktop for video.

Does more system RAM replace VRAM?

It prevents crashes and allows offloading, at a large speed cost. It is a supplement, not a substitute.

When is renting cheaper?

Below roughly 10 hours of generation a week, almost always. A 3090 draws 350 W; add the purchase price and the cloud subscriptions win for casual use and for video.

Which cloud for which job?

BitVector for a web app with the models preloaded; PirateDiffusion for the full model library, ComfyUI workflows and video from Telegram. Both are linked from every model page.

Links and sources

Models in this guide

Written by
Quartermaster

Related guides

Errors & fixes
CUDA out of memory: how much VRAM each AI model needs and how to fit it
A VRAM table for SD 1.5, SDXL, Flux, FLUX.2, Qwen-Image, Z-Image, Wan 2.2 and HunyuanVideo, and the techniques that make a model fit: fp8 and GGUF quantisation, CPU offload, tiled VAE, resolution and frame limits, and when to stop fighting and run it hosted.
By Captain
2026-05-18
ComfyUI
What is ComfyUI? A beginner's guide to nodes, workflows and the first graph
ComfyUI explained for people coming from Automatic1111, Fooocus or a web app: why it is a graph, the seven nodes in the default workflow and what each does, how to load a workflow from an image, where models go, the Manager for missing nodes, and when a simpler tool or the cloud is the better choice.
By Lookout
2025-12-17
Apps
Running Stable Diffusion, SDXL and Flux on a Mac (Apple Silicon): Draw Things, ComfyUI and what to expect
What an M1 to M4 Mac can and cannot do for AI images and video: the three ways to run models (Draw Things, ComfyUI through PyTorch MPS, DiffusionBee), memory tiers by Mac configuration, realistic speeds for SD 1.5, SDXL and Flux, the settings that avoid black images and crashes, and the models that are worth the wait on a Mac.
By Captain
2026-01-21
Models
Which AI image models to try first, and why: a ranked starting list
Seven models for a first week, chosen for forgiveness rather than hype: one SDXL photo fine-tune, one anime fine-tune, Flux dev, a fast turbo model, an inpainting model, a video model and an upscaler, each with the reason, the memory it needs and the prompt style it expects.
By Captain
2025-09-10
Video models
AI video generation for beginners: image-to-video, text-to-video and the first clip with Wan, LTX-2 and H3
A first week with open video models: the difference between text-to-video and image-to-video and why to start with the latter, the three model families worth learning (Wan 2.2, LTX-2, MiniMax H3), frame counts and resolutions that work, how to write a motion prompt, what to expect from a 16 or 24 GB card against the cloud, and how to turn five-second clips into something longer.
By Lookout
2026-04-29

More guides

Models
FLUX.1 Dev: how to prompt it, best settings and creative ideas
A practical guide to FLUX.1 Dev by Black Forest Labs: natural-language prompting, guidance and step settings, LoRA stacking, text rendering and prompt ideas that play to its strengths.
By Captain
2026-09-30
Models
Stable Diffusion XL 1.0: prompts, settings and what it still does best
How to get the most out of the official SDXL 1.0 base model: native resolutions, CFG and sampler settings, the refiner, negative prompts and ideas for styles where SDXL still shines.
By Captain
2026-10-01
Models
Stable Diffusion 1.5: the classic model, prompted properly
Stable Diffusion 1.5 is still worth knowing: the right resolution, CFG and sampler, how to use its huge library of LoRAs and embeddings, and creative prompt ideas suited to a 512-pixel model.
By Captain
2026-10-02
Models
FLUX.2 Dev: prompting the newest Black Forest Labs model
FLUX.2 Dev brings larger prompts, better text, stronger realism and multi-reference editing. This guide covers the turbo workflow, guidance and resolution settings, prompt structure and ideas that use its new strengths.
By Captain
2026-10-03
Models
Qwen-Image: long prompts, perfect text and bilingual posters
Qwen-Image is the model to reach for when the words in the picture matter. Learn how to write its long descriptive prompts, render English and Chinese text accurately, set CFG and steps, and explore layout-heavy creative ideas.
By Captain
2026-09-29
Models
SDXL-Lightning: four-step generation on any SDXL checkpoint
ByteDance's SDXL-Lightning LoRA turns a 30-step SDXL render into a 4-step one. Learn the step counts, the CFG you must use, the sampler, and how to combine it with your favorite checkpoints and style LoRAs.
By Captain
2026-09-30