Model
Trends
.ai
Model
Trends
.ai
the best open source ai models

ControlNet explained: pose, depth, edges and reference images for SD 1.5, SDXL and Flux

Topic: Prompting
By Captain
Published 2026-02-18
How ControlNet makes a model follow a pose, a depth map, an edge drawing or a scribble instead of guessing the composition from the prompt: the common control types and when to use each, the preprocessor step, strength and end-percent settings, the family-specific files, and the Flux-era alternatives (Flux Depth, Canny, Kontext).

Overview

A prompt says what to draw; ControlNet says where. It is a side network, trained per family, that reads a control image (a stick-figure pose, a depth map, an edge drawing, a scribble, a segmentation map) and steers the denoiser so the result matches that structure. The prompt still decides subject, style and light; the control image fixes the composition. It is how people get the same pose across ten characters, turn a 3D render into a photo, or keep a product's exact shape.
Two steps always happen. First a preprocessor turns your source photo into the control image (OpenPose draws the skeleton, MiDaS or Depth Anything estimates depth, Canny finds edges, Lineart traces outlines). Then the ControlNet model for your family consumes that image during generation. Feeding the raw photo instead of the preprocessed map is the number one beginner mistake; the second is using an
SD 1.5
ControlNet on an
SDXL
checkpoint
.
Flux
and the 2025 models changed the landscape: Black Forest Labs ships Flux Depth and Flux Canny as full models and
LoRAs
, Union ControlNets cover several types in one file for SDXL and Flux, and instruction editors (
Qwen Image
Edit, Kontext) handle many "keep this, change that" jobs without a control image. The catalog on this site tags ControlNets under "other" with their family; the model page names the control type.

Reference

Name
Type
What it is
OpenPose
control type
Body, hand and face keypoints as a stick figure. Pose transfer, choreography, consistent characters. Weak on hands unless the hand model is on.
Depth (MiDaS, Zoe, Depth Anything)
control type
Grey depth map. Keeps spatial layout and camera while allowing full restyling. The most forgiving control for photo to illustration.
Canny / Lineart / SoftEdge (HED)
control type
Edge drawings. Canny is strict and sharp, Lineart for drawings, SoftEdge loose. Product shapes, architecture, logos.
Scribble
control type
Rough strokes you draw yourself become a composition. Low strength (0.5-0.7) or the result looks like the scribble.
Tile
control type
Guides an upscale or a detail pass to stay faithful to the low-resolution input. The engine behind most "ultimate upscale" workflows.
IP-Adapter / Redux / reference
image conditioning
Not a ControlNet: conditions on an image's style or face rather than its structure. Combine with a ControlNet to get both.
Strength / weight
0.4-1.0
How hard the structure is enforced. 0.8 for pose and depth, 0.5-0.7 for edges and scribble.
Start / end percent
0.0-1.0
Which part of the denoising the control applies to. Ending at 0.6-0.8 lets the model finish details freely and removes the "traced" look.
Family files
SD 1.5: lllyasviel control_v11 | SDXL: diffusers / xinsir Union | Flux: Flux Depth/Canny, Shakker Union, InstantX
Each family needs its own ControlNet; sizes 0.7-2.5 GB. Union files bundle several types.
AD
Pose and depth control in the browser
BitVector Prism includes pose and depth control for its SDXL and Flux models: upload a reference, pick the control, type the prompt. No preprocessors to install.

Step by step

  1. Pick the control type for the job: pose for people, depth for scenes and restyling, Canny or Lineart for hard shapes, scribble for ideas, tile for
    upscaling
    .
  2. Download the ControlNet for your checkpoint's family into models/controlnet (
    ComfyUI
    ) or models/ControlNet (A1111,
    Forge
    ). Union files save disk if you use several types.
  3. Load your source image and run the matching preprocessor (ComfyUI: the controlnet_aux nodes; A1111: the ControlNet extension does it inline). Preview the control image; fix it if the skeleton missed a limb.
  4. ComfyUI: Apply ControlNet (Advanced) between CLIP Text Encode and KSampler, wire the control image and the ControlNet model, strength 0.8, start 0, end 0.8.
  5. A1111 / Forge: enable ControlNet unit 0, choose preprocessor and model of the same type, weight 0.8, ending control step 0.8, pixel perfect on.
  6. Write the prompt for what you want to see, not for the pose (the pose is handled). Generate four seeds.
  7. Lower strength or the end percent if the image looks traced or flat; raise them if the structure drifts. Stack a second unit (pose + depth) for stubborn compositions.
  8. On the cloud,
    PirateDiffusion
    applies ControlNet when you reply to a source image with /render and a control flag;
    BitVector
    offers pose and depth in its image tools.

Examples

Pose transfer in ComfyUI (SDXL)

Load Image (reference photo) -> DWPose Estimator (body + hands) -> Apply ControlNet Advanced (xinsir controlnet-union-sdxl, type openpose, strength 0.8, end 0.75) -> KSampler
Prompt: a knight in weathered plate armour standing on a cliff at dawn, cinematic, 35mm

Photo to watercolour with depth (A1111 / Forge)

ControlNet unit 0: preprocessor depth_anything, model control_sdxl_depth, weight 0.7, end 0.7, pixel perfect
Prompt: loose watercolour illustration of a harbour town, paper texture, soft morning light

Flux Canny LoRA

Load Diffusion Model (flux1-dev-fp8) -> LoraLoaderModelOnly (flux1-canny-dev-lora, 0.85) -> InstructPixToPixConditioning (canny image) -> KSampler (guidance 30)

Tips

  • Depth beats Canny for photo-to-illustration because it leaves texture free; Canny beats depth for logos and products because it keeps edges exact.
  • End percent is the setting most people never touch and the one that fixes the plastic, traced look.
  • Resize the control image to the generation size with the same aspect ratio; mismatched shapes crop the skeleton.
  • OpenPose from a photo of yourself is the fastest way to get a consistent pose library for a character.
  • For Flux, the Depth and Canny LoRAs from Black Forest Labs are lighter than the full control models and good enough for most jobs.
  • Instruction editors (Qwen Image Edit, Kontext) often replace a ControlNet pass for "same scene, different subject" edits; see the Qwen Image Edit guide.

Troubleshooting

ControlNet does nothing

Why it happens
Family mismatch (SD 1.5 ControlNet on SDXL), or the raw photo was fed instead of the preprocessed map.

How to fix it
Match the family; run the preprocessor and preview its output.

Result looks traced and flat

Why it happens
Strength 1.0 with end percent 1.0.

How to fix it
Strength 0.7, end 0.7; let the model finish on its own.

Skeleton misses hands or a limb

Why it happens
Preprocessor detected poorly on a cluttered photo.

How to fix it
Use DWPose with hands enabled, crop the source, or edit the skeleton in an OpenPose editor node.

Out of memory with two units

Why it happens
Each ControlNet is up to 2.5 GB.

How to fix it
Use a Union model for both types, or fp8 ControlNets; see the
VRAM
guide.

Flux ControlNet gives noise

Why it happens
Wrong conditioning node or
CFG
above 1 on a guidance-distilled model.

How to fix it
Use the node the model page specifies and CFG 1 with guidance per the page.

AD
PirateDiffusion
ControlNet by replying to a photo
On PirateDiffusion you reply to any image in Telegram with /render and a control flag and the right ControlNet for the model is applied server-side. Unlimited renders on a fixed price.

Questions

Is ControlNet a LoRA?

No. It is a separate network that reads an image; a LoRA changes the model's weights from text. Flux's Depth and Canny "LoRAs" are an exception: control models shipped in LoRA form.

Can I use ControlNet with video?

Yes:
Wan
VACE and the Wan Fun ControlNets take pose or depth sequences; the video guides on this site cover them.

Do I need the preprocessor models installed?

Yes, they download on first use (controlnet_aux in ComfyUI, the extension in A1111). They are small.

Which is better, IP-Adapter or ControlNet?

Different jobs: IP-Adapter copies look or identity from an image, ControlNet copies structure. Many workflows use both.

Does this work on the cloud?

PirateDiffusion and BitVector have the ControlNets for their models installed; you provide the source image and the control type.

Links and sources

Models in this guide

Written by
Captain

Related guides

Models
Checkpoint vs LoRA vs embedding vs ControlNet: which file does what
The six kinds of model files you will meet, what each one changes, how big it is, where it goes and when to use it: checkpoints and fine-tunes, LoRAs and their LyCORIS cousins, textual inversion embeddings, ControlNets and IP-Adapters, VAEs and upscalers.
By Quartermaster
2025-10-08
Prompting
img2img, inpainting and outpainting basics: fixing hands, swapping objects and extending a picture
The three ways to generate from an existing picture instead of from noise: img2img for restyling with a denoise value, inpainting for repainting only a masked area (the standard fix for hands and faces), and outpainting for extending the canvas; with the denoise, padding and model choices for SD 1.5, SDXL, Flux Fill and Qwen Image Edit.
By Captain
2025-12-03
Prompting
Qwen-Image-Edit guide: instruction prompts, multi-image editing and the LoRAs that make it better
How to get clean edits from Qwen-Image-Edit: writing instructions instead of descriptions, keeping identity and layout, multi-image inputs, text replacement, pose and depth control, the lightning LoRAs for 4-step edits and the anything-to-real and slider LoRAs trending in the Qwen2 family.
By Captain
2026-05-26
ComfyUI
What is ComfyUI? A beginner's guide to nodes, workflows and the first graph
ComfyUI explained for people coming from Automatic1111, Fooocus or a web app: why it is a graph, the seven nodes in the default workflow and what each does, how to load a workflow from an image, where models go, the Manager for missing nodes, and when a simpler tool or the cloud is the better choice.
By Lookout
2025-12-17
Models
Upscaling AI images: ESRGAN models, hires-fix, tiled diffusion and when each one is right
The three families of upscaling and what they are for: fast pixel upscalers (4x-UltraSharp, Real-ESRGAN, SwinIR, DAT) for clean enlargement, hires-fix for adding detail at generation time, and tiled diffusion (Ultimate SD Upscale, SUPIR, Flux-based) for turning a 1024 render into a detailed 4K print; with denoise values, tile sizes and the models worth downloading.
By Quartermaster
2026-03-04

More guides

Models
FLUX.1 Dev: how to prompt it, best settings and creative ideas
A practical guide to FLUX.1 Dev by Black Forest Labs: natural-language prompting, guidance and step settings, LoRA stacking, text rendering and prompt ideas that play to its strengths.
By Captain
2026-09-30
Models
Stable Diffusion XL 1.0: prompts, settings and what it still does best
How to get the most out of the official SDXL 1.0 base model: native resolutions, CFG and sampler settings, the refiner, negative prompts and ideas for styles where SDXL still shines.
By Captain
2026-10-01
Models
Stable Diffusion 1.5: the classic model, prompted properly
Stable Diffusion 1.5 is still worth knowing: the right resolution, CFG and sampler, how to use its huge library of LoRAs and embeddings, and creative prompt ideas suited to a 512-pixel model.
By Captain
2026-10-02
Models
FLUX.2 Dev: prompting the newest Black Forest Labs model
FLUX.2 Dev brings larger prompts, better text, stronger realism and multi-reference editing. This guide covers the turbo workflow, guidance and resolution settings, prompt structure and ideas that use its new strengths.
By Captain
2026-10-03
Models
Qwen-Image: long prompts, perfect text and bilingual posters
Qwen-Image is the model to reach for when the words in the picture matter. Learn how to write its long descriptive prompts, render English and Chinese text accurately, set CFG and steps, and explore layout-heavy creative ideas.
By Captain
2026-09-29
Models
SDXL-Lightning: four-step generation on any SDXL checkpoint
ByteDance's SDXL-Lightning LoRA turns a 30-step SDXL render into a 4-step one. Learn the step counts, the CFG you must use, the sampler, and how to combine it with your favorite checkpoints and style LoRAs.
By Captain
2026-09-30