Model
Trends
.ai
Model
Trends
.ai
the best open source ai models

Unet Loader (GGUF) node in ComfyUI

Unet Loader (GGUF) loads quantised Q8, Q6, Q5 and Q4 .gguf versions of Flux, Wan, Hunyuan, Qwen-Image and other large transformers so they run on 6-12 GB GPUs; DualCLIPLoader (GGUF) does the same for T5 text encoders.
Node pack: ComfyUI-GGUF
Category: Loaders

What Unet Loader (GGUF) does

GGUF is the quantisation format from llama.cpp, brought to image and video models by city96's ComfyUI-GGUF pack. A Flux Dev transformer that needs 24 GB in bf16 becomes 12 GB at Q8, 7 GB at Q4_K_S, with quality that degrades gracefully. Unet Loader (GGUF) reads these files from models/unet (or models/diffusion_models) and outputs a normal MODEL; weights are dequantised on the fly, so it is a little slower than fp8 but fits where fp8 does not. DualCLIPLoader (GGUF) and CLIPLoader (GGUF) load t5xxl and umt5 GGUF text encoders in the same way.
LoRAs, ControlNets and IP-Adapters work with GGUF models. The rest of the graph is unchanged: text encoders, VAE, sampler and ModelSampling nodes stay the same as for the safetensors version.

Inputs

Name
Type
What it is
unet_name
COMBO
A .gguf file from models/unet or models/diffusion_models.

Outputs

Name
Type
What it is
MODEL
MODEL
The quantised model for LoRA loaders and KSampler.
AD
Skip the setup: ComfyUI workflows preinstalled on the cloud
BitVector runs ready-made ComfyUI workflows on its own GPUs. No install, no missing nodes, no red boxes. Open it on your phone or laptop and generate in a minute.

How to use Unet Loader (GGUF)

  1. Install ComfyUI-GGUF from Manager (it needs the gguf Python package, installed automatically).
  2. Download the Q8_0 (best) or Q4_K_S (smallest) GGUF of your model into models/unet.
  3. Replace Load Diffusion Model with Unet Loader (GGUF) and pick the file.
  4. Optionally use DualCLIPLoader (GGUF) with t5-v1_1-xxl-encoder-Q5_K_M.gguf to shrink the text encoder.
  5. Keep the rest of the graph; render as usual.

Settings and tips

  • Q8_0 is visually identical to fp16 for most models; Q5_K_M is the sweet spot for 8 GB cards; Q4 shows small texture loss.
  • GGUF decoding is CPU/GPU work: expect 10-30% slower steps than fp8 on cards that support fp8.
  • For Wan 2.2 there are GGUFs of both the high and the low noise models.
  • Do not mix: the VAE and CLIP-L stay as safetensors files.
  • --lowvram together with GGUF Q4 lets Flux run on 6 GB with slow but working results.

Troubleshooting Unet Loader (GGUF)

The .gguf file does not appear in the dropdown

Why it happens
The pack is not installed (the node itself would be red), or the file is in the wrong folder.

How to fix it
Install ComfyUI-GGUF, put the file in models/unet or models/diffusion_models and press R.

Error: unknown GGUF architecture / unsupported quant type

Why it happens
The GGUF is a language model, or a quantisation type newer than the installed pack.

How to fix it
Download an image-model GGUF from the city96 or QuantStack repositories and update ComfyUI-GGUF.

LoRA produces warnings about missing keys

Why it happens
The LoRA targets layers in the original naming; most are mapped, but some trainers produce unusual key names.

How to fix it
Try the LoRA with the safetensors model to confirm it works; update the GGUF pack, which improves key mapping over time.

Still out of memory at Q4

Why it happens
The text encoder (T5 at 10 GB fp16) is the real consumer.

How to fix it
Use the GGUF or fp8 T5 and the fp8 CLIP; add --lowvram.

AD
Run this workflow from your phone
Every workflow on this page is preinstalled on BitVector with the models and custom nodes already in place. Pick one, type a prompt, done. Zero setup, nothing to download.

Questions about Unet Loader (GGUF)

GGUF or fp8?

fp8 is faster on RTX 40/50 cards when the model fits; GGUF Q8 has slightly better quality than fp8 and Q4-Q6 go where fp8 cannot.

Can I make my own GGUF?

Yes, with the convert script in the ComfyUI-GGUF repository and llama.cpp quantise tools; most people download ready files.

Related nodes

Load Diffusion Model
ComfyUI core
Load Diffusion Model (UNETLoader) loads a bare denoising network such as Flux, SD3.5, Wan or HunyuanVideo from models/diffusion_models, with a weight_dtype option for fp8 to save VRAM.
DualCLIPLoader
ComfyUI core
DualCLIPLoader loads two text encoders at once, for example clip_l plus t5xxl for Flux or clip_g plus t5xxl for SD3, and outputs one CLIP object for the prompt nodes.
Load LoRA
ComfyUI core
Load LoRA (LoraLoader) applies a LoRA file to the MODEL and CLIP with separate strengths, so a style, character or concept can be added to any checkpoint without merging.
ModelSamplingFlux, ModelSamplingSD3 and ModelSamplingAuraFlow
ComfyUI core
The ModelSampling nodes set the noise schedule shift for flow-matching models (Flux, SD3.5, Wan, Hunyuan, Qwen-Image, Chroma, Lumina) so samplers spend the right amount of time on composition versus detail.
ComfyUI Manager
ComfyUI-Manager
ComfyUI Manager is the extension that installs, updates and fixes custom node packs and models from inside the interface, resolves missing nodes in imported workflows and snapshots your setup.
AD
Tired of fixing nodes? Let the cloud do it
BitVector keeps hundreds of ComfyUI workflows installed, updated and tested on fast cloud GPUs. No Python, no CUDA errors, no VRAM limits. Works from any browser.

More ComfyUI nodes

CLIP Text Encode (Prompt)
ComfyUI core
CLIP Text Encode turns a text prompt into CONDITIONING using the model text encoder. One node holds the positive prompt, a second one the negative prompt.
KSampler
ComfyUI core
KSampler runs the denoising loop: it takes the model, prompts and a latent and produces the finished latent image, controlled by seed, steps, cfg, sampler, scheduler and denoise.
Load Checkpoint
ComfyUI core
Load Checkpoint (CheckpointLoaderSimple) opens a .safetensors or .ckpt model file and hands out the three parts every workflow needs: the diffusion MODEL, the CLIP text encoder and the VAE.
Save Image
ComfyUI core
Save Image writes the IMAGE tensor to ComfyUI/output as a PNG, with the whole workflow embedded in the file metadata so the picture can be dragged back into ComfyUI to restore the graph.
VAE Decode
ComfyUI core
VAE Decode converts the sampled LATENT into a pixel IMAGE with the VAE; VAE Decode (Tiled) does the same in tiles for very large images.
Empty Latent Image
ComfyUI core
Empty Latent Image creates the blank latent canvas (width, height, batch_size) that text-to-image sampling starts from; dimensions must be multiples of 8 and match the model family.