Model
Trends
.ai
Model
Trends
.ai
en iyi açık kaynak yapay zeka modelleri

Checkpoint vs LoRA vs embedding vs ControlNet: which file does what

Konu: Modeller
Quartermaster tarafından
2025-10-08 tarihinde yayımlandı
The six kinds of model files you will meet, what each one changes, how big it is, where it goes and when to use it: checkpoints and fine-tunes, LoRAs and their LyCORIS cousins, textual inversion embeddings, ControlNets and IP-Adapters, VAEs and upscalers.
Bu sayfa henüz çevrilmedi, bu yüzden İngilizce gösteriliyor.

Genel bakış

A new user downloads a "model" and discovers it is one of six different things that go in six different folders and do six different jobs. The catalog on this site labels every entry with its type (base model,
LoRA
, embedding, other) and its family for that reason. This guide explains the types once, so the labels make sense.
The rule of thumb: a checkpoint is the whole painter; a LoRA teaches the painter a new subject or style; an embedding teaches the painter a new word; a
ControlNet
or IP-Adapter tells the painter where things go or what to copy from a reference; the VAE is the painter's eyes converting between the canvas and the compressed working space; an
upscaler
is a separate craftsman who enlarges the finished picture. Only the checkpoint can work alone. Everything else needs a checkpoint of the matching family.
Sizes tell you a lot. 2 to 7 GB for an
SDXL
checkpoint and 12 to 24 GB for
Flux
; 10 to 700 MB for a LoRA; under 1 MB for an embedding; 1 to 2.5 GB for a ControlNet; 300 MB for a VAE; 60 MB for an ESRGAN upscaler. If a file called a LoRA is 6 GB, it is a checkpoint. The model page here shows the size and the hash so you can tell before downloading.

Başvuru

Ad
Tür
Ne olduğu
Checkpoint (base model)
2-24 GB
The full denoiser, usually with text encoder and VAE bundled. Fine-tunes (DreamShaper, Juggernaut,
Illustrious
, Pony) are checkpoints retrained on new data. Folder: models/checkpoints (
ComfyUI
), models/Stable-diffusion (A1111,
Forge
). Flux,
Qwen
and
Wan
ship as separate diffusion_models + text_encoders + vae files.
LoRA
10-700 MB
A small weight difference applied on top of a checkpoint. Characters, styles, concepts, sliders, motion. Needs a trigger word and a weight. Folder: models/loras or models/Lora.
LyCORIS / LoCon / LoHa / DoRA
10-700 MB
LoRA variants that touch more layers or use a different factorisation. Loaded like LoRAs in every current UI.
Embedding (textual inversion)
10-200 KB
A learned token for the text encoder. Called by its file name in the prompt. Common as negative embeddings (bad-hands, easynegative) on
SD 1.5
and SDXL; rare on Flux. Folder: models/embeddings.
ControlNet / T2I-Adapter
0.3-2.5 GB
A side network that conditions the denoiser on a pose, depth map, edge map or scribble. Family specific. Folder: models/controlnet.
IP-Adapter / reference
50 MB-1 GB
Conditions on an image instead of text: style transfer, face reference. Needs a matching CLIP vision model.
VAE
160-330 MB
Encoder/decoder between pixels and latents. Usually bundled in checkpoints; separate files fix washed-out colours (SDXL fp16-fix) or come with new families (Flux ae.safetensors, Wan VAE).
Upscaler
2-70 MB
ESRGAN, Real-ESRGAN, SwinIR, DAT: separate networks that enlarge pixels after generation. Folder: models/upscale_models.
AD
No folders, no file types, no VAE hunting
BitVector Prism preloads the checkpoint, the LoRAs, the VAE and the ControlNets that belong together. Pick a model from this site, pick an add-on and generate in the browser.

Adım adım

  1. Identify the file type from the catalog label and the size before downloading. Base model, LoRA, embedding or other on the card; the exact bytes on the model page.
  2. Match the family. Every non-checkpoint file names the family it was built for (SDXL, Flux,
    Wan 2.2
    14B); it only works with checkpoints of that family.
  3. Put the file in the folder for its type (table above) and refresh the UI.
  4. Checkpoint: select it as the model. For split families (Flux, Qwen, Wan,
    Z-Image
    ) load the diffusion model, the text encoders and the VAE separately in ComfyUI.
  5. LoRA: Load LoRA node after the checkpoint, or <lora:name:0.8> in the prompt, plus the trigger word.
  6. Embedding: type the file name in the prompt (or the
    negative prompt
    for negative embeddings); ComfyUI uses embedding:name.
  7. ControlNet: Apply ControlNet node with the preprocessed image (pose, depth); start with strength 0.8 and end percent 0.8.
  8. VAE and upscaler: select the VAE in the loader or settings when colours are wrong; run the upscaler as a final step through Upscale Image (using Model).

Örnekler

A full SDXL stack in a Forge prompt

Checkpoint: juggernautXL_v9.safetensors | VAE: sdxl_vae_fp16fix.safetensors
Prompt: zavy cinematic, <lora:zavy-cinematic-xl:0.7> a knight resting by a campfire, night, embers, 35mm
Negative: easynegative, bad-hands-5, blurry | ControlNet: openpose, weight 0.8 | Upscaler: 4x-UltraSharp 2x

The same stack as ComfyUI nodes

Load Checkpoint -> Load LoRA -> Apply ControlNet (OpenPose image) -> KSampler -> VAE Decode (fp16-fix VAE) -> Upscale Image (using Model: 4x-UltraSharp)

İpuçları

  • The catalog type "other" covers ControlNets, VAEs,
    upscalers
    and workflows; the model page says which.
  • Merges are checkpoints made by averaging other checkpoints; they inherit the family of their parents and keep LoRA compatibility.
  • An
    inpainting
    checkpoint is a checkpoint with extra input channels; it needs the inpainting workflow and is not a drop-in for text-to-image.
  • Negative embeddings were the SD 1.5 way of fixing hands; on SDXL they help less and on Flux they do nothing because the model has no
    CFG
    negative.
  • Flux Fill, Flux Redux, Flux Depth and Canny are checkpoints or adapters, not LoRAs, even though they ship beside the LoRAs on the hub.
  • When in doubt, the cloud partners have the right combination preloaded: pick the model and the LoRA and the matching VAE and ControlNet are handled for you.

Sorun giderme

Downloaded a "LoRA" and it is 6 GB

Neden olur
It is a checkpoint; the uploader labelled it loosely.

Nasıl düzeltilir
Put it in the checkpoints folder and select it as the model. The catalog label here is reliable.

Embedding does nothing

Neden olur
Wrong family (SD 1.5 embedding on SDXL), or the name typed does not match the file name.

Nasıl düzeltilir
Match the family; type the exact file name without extension; in ComfyUI use embedding:name.

ControlNet gives noise or ignores the pose

Neden olur
ControlNet for another family, or the raw photo was fed instead of the preprocessed map.

Nasıl düzeltilir
Use the ControlNet built for the checkpoint family and run the preprocessor node first.

Colours look washed out

Neden olur
Missing or wrong VAE.

Nasıl düzeltilir
Load the family's VAE; on SDXL the fp16-fix VAE.

Model loads but generates black images

Neden olur
fp16 overflow in the VAE or a checkpoint that needs --no-half-vae.

Nasıl düzeltilir
Use the fp16-fix VAE or the no-half-vae flag; on Flux check that the text encoder precision matches the build.

AD
PirateDiffusion
Checkpoints, LoRAs, embeddings and ControlNets, already wired
PirateDiffusion keeps thousands of model files organised for you on its own GPUs; you name the checkpoint and the LoRA tags in a Telegram command and the matching VAE and ControlNets are applied. Unlimited generation, fixed price.

Sorular

Which one should a beginner download first?

One checkpoint of the family you want to use. LoRAs come second, once you know which base you like. Embeddings and ControlNets come when a specific problem needs them.

Can I merge a LoRA into a checkpoint?

Yes, every UI has a merge tool, and the result is a new checkpoint. It saves a node but freezes the weight; most people keep LoRAs separate.

Is a fine-tune the same as a LoRA?

No. A fine-tune retrains the whole checkpoint (gigabytes); a LoRA is a small add-on (megabytes). Many fine-tunes started as LoRAs merged in.

Do Flux and Qwen use embeddings?

Practically never. Their LLM text encoders are not trained that way; style and concept work is done with LoRAs.

What is a workflow file?

A ComfyUI JSON that wires these components together. The catalog lists some as "other"; they contain no weights, only the recipe.

Bağlantılar ve kaynaklar

Bu rehberdeki modeller

Yazan
Quartermaster

İlgili rehberler

Modeller
What is a LoRA and how to use one in ComfyUI, Forge, Automatic1111 and on the cloud
LoRAs explained for people who want results, not maths: what a LoRA file changes, why it must match the base model's family, trigger words, weights, stacking several LoRAs, the folder and syntax for each UI, and how to try a LoRA on the cloud before downloading anything.
Captain tarafından
2026-05-20
Uygulamalar
How to install a LoRA: folders, file checks and first use in ComfyUI, Forge, Automatic1111 and Draw Things
The practical installation walkthrough: where to download a LoRA from, how to check the file hash against the model page, which folder it goes in for every popular UI, how to refresh the list, how to load it the first time, and the three ways to use one on the cloud without installing anything.
Lookout tarafından
2025-09-24
Modeller
How generative AI image models work: diffusion, latents and text encoders explained simply
The mechanics behind Stable Diffusion, SDXL, Flux and the video models without the maths: what noise has to do with it, why models work in a compressed latent space, what the text encoder and VAE do, what steps and guidance really change, and why training data decides what a model can draw.
Quartermaster tarafından
2025-08-28
Prompt yazımı
ControlNet explained: pose, depth, edges and reference images for SD 1.5, SDXL and Flux
How ControlNet makes a model follow a pose, a depth map, an edge drawing or a scribble instead of guessing the composition from the prompt: the common control types and when to use each, the preprocessor step, strength and end-percent settings, the family-specific files, and the Flux-era alternatives (Flux Depth, Canny, Kontext).
Captain tarafından
2026-02-18
Modeller
Upscaling AI images: ESRGAN models, hires-fix, tiled diffusion and when each one is right
The three families of upscaling and what they are for: fast pixel upscalers (4x-UltraSharp, Real-ESRGAN, SwinIR, DAT) for clean enlargement, hires-fix for adding detail at generation time, and tiled diffusion (Ultimate SD Upscale, SUPIR, Flux-based) for turning a 1024 render into a detailed 4K print; with denoise values, tile sizes and the models worth downloading.
Quartermaster tarafından
2026-03-04

Diğer rehberler

Modeller
FLUX.1 Dev: how to prompt it, best settings and creative ideas
A practical guide to FLUX.1 Dev by Black Forest Labs: natural-language prompting, guidance and step settings, LoRA stacking, text rendering and prompt ideas that play to its strengths.
Captain tarafından
2026-09-30
Modeller
Stable Diffusion XL 1.0: prompts, settings and what it still does best
How to get the most out of the official SDXL 1.0 base model: native resolutions, CFG and sampler settings, the refiner, negative prompts and ideas for styles where SDXL still shines.
Captain tarafından
2026-10-01
Modeller
Stable Diffusion 1.5: the classic model, prompted properly
Stable Diffusion 1.5 is still worth knowing: the right resolution, CFG and sampler, how to use its huge library of LoRAs and embeddings, and creative prompt ideas suited to a 512-pixel model.
Captain tarafından
2026-10-02
Modeller
FLUX.2 Dev: prompting the newest Black Forest Labs model
FLUX.2 Dev brings larger prompts, better text, stronger realism and multi-reference editing. This guide covers the turbo workflow, guidance and resolution settings, prompt structure and ideas that use its new strengths.
Captain tarafından
2026-10-03
Modeller
Qwen-Image: long prompts, perfect text and bilingual posters
Qwen-Image is the model to reach for when the words in the picture matter. Learn how to write its long descriptive prompts, render English and Chinese text accurately, set CFG and steps, and explore layout-heavy creative ideas.
Captain tarafından
2026-09-29
Modeller
SDXL-Lightning: four-step generation on any SDXL checkpoint
ByteDance's SDXL-Lightning LoRA turns a 30-step SDXL render into a 4-step one. Learn the step counts, the CFG you must use, the sampler, and how to combine it with your favorite checkpoints and style LoRAs.
Captain tarafından
2026-09-30