How generative AI image models work: diffusion, latents and text encoders explained simply
Topic: Models
By Quartermaster
Published 2025-08-28
The mechanics behind Stable Diffusion, SDXL, Flux and the video models without the maths: what noise has to do with it, why models work in a compressed latent space, what the text encoder and VAE do, what steps and guidance really change, and why training data decides what a model can draw.
Overview
Every current image model is a denoiser. During training it is shown millions of images with increasing amounts of random noise added, together with the caption of each image, and it learns to predict what noise was added. At generation time the process runs backwards: start from pure noise, ask the model what noise is in there given your prompt, remove a little of it, repeat twenty or thirty times, and an image that matches the caption emerges. That is the whole trick; "diffusion" is the name of this gradual add-noise, remove-noise process.
Working on full-size pixels would be far too slow, so since Stable Diffusion 1.x the denoising happens in a compressed space called the latent. A separate small network, the VAE, squeezes a 1024x1024 image into a 128x128 grid of numbers and expands it back at the end. The denoiser (a U-Net in
SD 1.5
and
SDXL
, a transformer in
Flux
,
Qwen
,
Z-Image
and the video models) never sees pixels. This is why image sizes need to be multiples of 8 or 16, and why the VAE decode step at the end can run
out of memory
on its own.
The prompt enters through a text encoder, which turns words into a list of numbers the denoiser can condition on. SD 1.5 uses one CLIP encoder with a 77-token limit, SDXL uses two, Flux adds a large T5 language model, and the 2025 and 2026 models (
Qwen Image
,
FLUX.2
, Z-Image,
Wan 2.2
) use full LLMs as encoders. That single change explains most of the difference in how prompts are written: tag lists for CLIP-based models, natural sentences for LLM-based ones. A
LoRA
, finally, is a small set of corrections to the denoiser weights (and sometimes the encoder) that steers what it draws; it does not add new knowledge from nowhere, it re-weights what the base model already learned.
Reference
Name | Type | What it is |
|---|---|---|
Denoiser (U-Net / DiT) | component | |
VAE | component | Compresses images to latents and back. A wrong or missing VAE gives washed-out or purple images. Flux and newer families ship their own 16-channel VAEs. |
Text encoder | component | CLIP-L, CLIP-G, T5-XXL or an LLM (Qwen2.5-VL, Mistral, Gemma). Decides how the prompt is understood and how long it can be. |
Scheduler / sampler | setting | The schedule of how much noise to remove at each step and the numerical method used. Changes texture and convergence speed, not content. |
Steps | setting | How many denoising iterations. Distilled models (Turbo, Lightning, Schnell, Klein) are trained to need 1 to 8. |
CFG / guidance | setting | Classifier-free guidance: how much the prompt-conditioned prediction is pushed away from the unconditioned one. Too high burns colours; guidance-distilled models ( Flux dev ) take a single guidance value instead. |
Latent size | setting | Image size divided by 8 (SD) or 16 (Flux, Wan). Models draw best near the resolution they were trained at. |
Flow matching | training method | The 2024+ replacement for classic diffusion noise schedules (Flux, SD3, Wan, Z-Image). Same idea, straighter path from noise to image, fewer steps needed. |
AD

See the theory in practice, no setup
BitVector Prism lets you switch between SDXL, Flux and Qwen models in one web app, so you can run the same prompt through three architectures and see the differences described here in a minute.
Step by step
- Noise in: thesamplercreates a latent grid of random numbers from the seed. Change the seed and everything downstream changes.
- Prompt in: the text encoder turns your words intoembeddings. On CLIP models anything past 77 tokens is chunked or dropped; on LLM-based models a paragraph is fine.
- Each step: the denoiser is run twice on SD and SDXL (with and without the prompt); the difference, scaled by CFG, points the latent toward the prompt. Flux dev runs once with a guidance value baked in.
- The sampler subtracts the predicted noise according to the schedule (Karras, exponential, simple, beta). After the last step the latent is a clean image in compressed form.
- Decode: the VAE expands the latent to pixels. Tiled VAE decode exists because this step alone can exceed the memory of an 8 GB card at large sizes.
- Optional second passes: hires-fix re-noises the result a little and denoises at a larger size;img2imgstarts from an existing image instead of pure noise; inpainting restricts the change to a mask.
- Video models add a time axis to the latent (a 5-second clip is one 3D latent) and are otherwise the same loop, which is why they need ten times the memory.
Examples
The same prompt through three architectures
SD 1.5 (CLIP, 77 tokens): "portrait of an old sailor, grey beard, wool cap, harbour background, film photo"
SDXL (two CLIPs): same words, 1024x1024, CFG 6, 28 steps
Flux dev (CLIP-L + T5): "A weathered old sailor with a grey beard and a navy wool cap stands at a harbour railing at dusk; shallow depth of field, Kodak Portra look." guidance 3.5, 24 steps
ComfyUI nodes mapped to the components
Load Checkpoint -> MODEL (denoiser), CLIP (text encoder), VAE
CLIP Text Encode (prompt) -> KSampler (sampler, steps, CFG, seed) -> VAE Decode -> Save Image
Tips
- A model cannot draw what was not in its training data. That is why anime fine-tunes know thousands of characters and photo models do not, and why a LoRA exists for almost every gap.
- Fine-tunes (DreamShaper, Juggernaut,Illustrious, Pony) are the same architecture as their base with different weights; a LoRA for the base usually works on the fine-tune, a LoRA for a very different fine-tune (Pony) often does not.
- Distilled models trade flexibility for speed. SDXL Lightning at 4 steps is fine for drafts; the full model at 28 steps is better for finals.
- The text encoder is often the biggest file. Flux with T5-XXL fp16 needs 9 GB just for the encoder, which is why fp8 and GGUF versions exist.
- Guidance distillation (Flux dev) and CFG are not the same thing. Using a CFG of 7 on Flux dev gives burnt images; use the model page settings.
- Everything here runs on the cloud partners exactly as locally; the difference is only who owns the GPU.
Troubleshooting
Images come out grey, washed out or with a purple tint
Why it happens
Missing or mismatched VAE, or an fp16 VAE overflow on SDXL.
How to fix it
Load the VAE that belongs to the family; on SDXL use the fp16-fix VAE.
Long prompt is partly ignored
Why it happens
CLIP token limit on SD 1.5 and SDXL.
How to fix it
Put the important words first, or move to a Flux, Qwen or Z-Image model with an LLM encoder.
Colours are burnt and contrast is extreme
Why it happens
CFG too high for the model, or CFG used on a guidance-distilled model.
How to fix it
Lower CFG to 5 to 7 on SDXL; use guidance 3 to 4 and CFG 1 on Flux dev.
Tiling or doubled subjects at large sizes
Why it happens
Latent much larger than the training resolution; the model repeats itself.
How to fix it
Generate at native size and upscale, or use hires-fix with 0.4 denoise.
Out of memory only at the very end
Why it happens
The VAE decode step at full size.
How to fix it
AD

Every architecture on one GPU cluster
PirateDiffusion hosts SD 1.5, SDXL, Flux, FLUX.2, Qwen, Z-Image and the video models side by side on Telegram, with unlimited generation on a fixed price. No VRAM maths, no VAE files to find.
Questions
Is the model copying images from its training set?
It stores weights, not images, and cannot retrieve a training picture. Near-duplicates of very frequent training images can appear (famous paintings, logos), which is a reason to avoid prompting for them.
Why do new models need more memory if the idea is the same?
Bigger denoisers (12 billion parameters for Flux against 0.9 for SD 1.5), bigger text encoders and larger latents. Quantised fp8 and GGUF versions bring them back down.
What is the difference between a checkpoint and a model?
A checkpoint is the saved file of a model (denoiser, often with encoder and VAE bundled). People use the words interchangeably.
Why does the same seed give a different image on another PC?
Different GPU, precision (fp16 vs bf16) or sampler implementation; the noise is the same but the arithmetic differs slightly. Same machine, same settings, same picture.
Does a higher step count always help?
No. Most samplers converge by 25 to 30 steps; beyond that you spend time for nothing. Distilled models get worse with more steps.
Links and sources
- Denoising Diffusion Probabilistic Models (Ho et al., 2020)
- High-Resolution Image Synthesis with Latent Diffusion Models (Rombach et al., 2022)
- Flow Matching for Generative Modeling (Lipman et al., 2022)
- Flux family page
- SD 1.5 family page
- BitVector web app
Models in this guide
Stable Diffusion 1.5 Models
Stable Diffusion XL 1.0 Models
Flux Models
Qwen Models
Zimage Models
Wan Models
Written by
Quartermaster
Related guides
Cloud services
Getting started with AI image generation: your first picture in ten minutes
A plain-language start for complete beginners: what a model is, the three ways to run one (web app, chat bot, your own PC), how to write a first prompt, what the settings mean and how to tell a good result from a lucky one.
By Captain
2025-08-14
Prompting
Steps, CFG, samplers, schedulers and seeds: the generation settings explained with safe defaults
What each setting in the generation panel changes, the values that work per model family (SD 1.5, SDXL, Flux, Qwen, Z-Image, Wan), which samplers are worth knowing, why distilled models break the usual rules, and how to use seeds to change one thing at a time.
By Quartermaster
2025-11-19
Models
Checkpoint vs LoRA vs embedding vs ControlNet: which file does what
The six kinds of model files you will meet, what each one changes, how big it is, where it goes and when to use it: checkpoints and fine-tunes, LoRAs and their LyCORIS cousins, textual inversion embeddings, ControlNets and IP-Adapters, VAEs and upscalers.
By Quartermaster
2025-10-08
Models
Flux vs SDXL vs SD 1.5 vs Qwen vs Z-Image: which model family to use in 2026
A practical comparison of the open image model families: prompt following, realism, anime and illustration, text rendering, speed, VRAM, LoRA ecosystem and licence, with a recommendation for each kind of work and hardware.
By Captain
2026-05-22
More guides
Models
FLUX.1 Dev: how to prompt it, best settings and creative ideas
A practical guide to FLUX.1 Dev by Black Forest Labs: natural-language prompting, guidance and step settings, LoRA stacking, text rendering and prompt ideas that play to its strengths.
By Captain
2026-09-30
Models
Stable Diffusion XL 1.0: prompts, settings and what it still does best
How to get the most out of the official SDXL 1.0 base model: native resolutions, CFG and sampler settings, the refiner, negative prompts and ideas for styles where SDXL still shines.
By Captain
2026-10-01
Models
Stable Diffusion 1.5: the classic model, prompted properly
Stable Diffusion 1.5 is still worth knowing: the right resolution, CFG and sampler, how to use its huge library of LoRAs and embeddings, and creative prompt ideas suited to a 512-pixel model.
By Captain
2026-10-02
Models
FLUX.2 Dev: prompting the newest Black Forest Labs model
FLUX.2 Dev brings larger prompts, better text, stronger realism and multi-reference editing. This guide covers the turbo workflow, guidance and resolution settings, prompt structure and ideas that use its new strengths.
By Captain
2026-10-03
Models
Qwen-Image: long prompts, perfect text and bilingual posters
Qwen-Image is the model to reach for when the words in the picture matter. Learn how to write its long descriptive prompts, render English and Chinese text accurately, set CFG and steps, and explore layout-heavy creative ideas.
By Captain
2026-09-29
Models
SDXL-Lightning: four-step generation on any SDXL checkpoint
ByteDance's SDXL-Lightning LoRA turns a 30-step SDXL render into a 4-step one. Learn the step counts, the CFG you must use, the sampler, and how to combine it with your favorite checkpoints and style LoRAs.
By Captain
2026-09-30

Model
Trends
.ai
© ModelTrends.ai
|
Made in Japan
|
© 2026