ControlNet: Adding Conditional Control to Text-to-Image Diffusion Models
ControlNet: Adding Conditional Control to Text-to-Image Diffusion Models

0. Abstract
This paper introduces ControlNet, which controls a pretrained diffusion model such as Stable Diffusion using additional inputs: edge maps, segmentation maps, keypoints, and more. It learns task-specific conditions and trains robustly even on datasets smaller than 50,000 examples.
1. Introduction
After writing text to generate an image, we might ask:
- Does text-prompt-based control adequately meet our need to generate the image we imagined?
- What framework should we build to accommodate a wide range of task conditions and user controls, including inputs other than text?
- Can we preserve a large model's capabilities learned from many images when adapting it to a specific task?
To answer these questions, the authors surveyed image-processing applications and made three observations, all pointing toward the need for less data and computation.
Three observations
- Task-specific datasets are not always as large as general text-to-image datasets.
- Most task-specific datasets contain fewer than 100,000 examples— smaller than LAION-5B.
- Large computational resources are not always available for data-driven image-processing solutions.
- Fast training methods are therefore important when optimizing large models, ideally allowing training on personal devices.
- Image-processing problems take many forms in their definitions, user controls, and image annotations.
- Tasks such as depth-to-image and pose-to-human require object- or scene-level interpretation of raw inputs. Building a procedural method for each one—constraining denoising, modifying multi-head attention activations, and so on—is difficult. End-to-end learning is essential for learned solutions across many tasks.
ControlNet is an end-to-end neural architecture that controls large image diffusion models such as Stable Diffusion with specific input conditions. It duplicates model weights into a trainable copy and a locked copy. The trainable copy learns conditional control from task-specific data, while the locked copy preserves the original capabilities. Special layers called zero convolutions connect the two blocks. Preserving production-ready weights makes training robust across dataset sizes. Because zero convolutions introduce no new noise, training is also as fast as fine-tuning the diffusion model.
Effects of a 1×1 convolution
- Reducing trainable parameters by changing channel counts
- Adding nonlinearity
- Computing object and positional information independently
3. Method
3.1. ControlNet
ControlNet manipulates the input conditions of neural network blocks to control the overall network's behavior. Such blocks include ResNet blocks, conv–BN–ReLU blocks, multi-head attention blocks, and transformer blocks.
For 2D features, given a feature map , a neural network block , parameterized by , transforms into another feature map .

All parameters in are locked and cloned into a trainable copy . The copied is trained using an external condition vector . The paper calls these the locked copy and trainable copy. Rather than directly updating the original weights, this preserves the quality learned from hundreds of millions of images and avoids overfitting on small datasets.
The blocks are connected by a special convolution layer , called a zero convolution: a 1×1 convolution whose weights and biases are initialized to zero. uses two parameter instances , producing the ControlNet structure below.
A trained ControlNet recognizes the semantic content of its input conditions. In other words, it reflects the meaning of depth, normals, edges, and other conditions in generation.

Because zero-convolution weights and biases start at zero, the first training step gives the following equations.

We can interpret them as follows.
These equations show that adding ControlNet does not affect the original network block before optimization. Its capabilities, functionality, and output quality remain intact. Subsequent optimization is as fast as fine-tuning, and much faster than training these layers from scratch.
Let us briefly examine gradients in a zero convolution. For a 1×1 convolution with weight and bias , spatial position denoted by , and channel index , the forward pass for an input map is as follows.
Before optimization, zero convolution has and . Wherever is nonzero, its gradients are:

Although zero convolution can make the gradient for the feature term zero, the gradients for weight and bias are unaffected. At the first gradient-descent step, the weight becomes a nonzero matrix as long as the feature is nonzero. Simply put, even if are initialized to zero, only needs to be nonzero for learning to occur.
Thus, zero convolution becomes a distinctive connection layer that evolves from zero into optimized parameters.
3.2. ControlNet in Image Diffusion Model
Stable Diffusion is used as the example.

To stabilize training, Stable Diffusion preprocesses its 512×512 images into 64×64 latent images, similarly to VQ-GAN. ControlNet must likewise encode image-based conditions into a 64×64 feature space to match convolution sizes. The paper uses a small network , with 4×4 kernels and 2×2 strides, to encode an image-space condition into a feature map .
The encoder transforms a 512×512 condition into a 64×64 feature map.
3.3. Training
Given a time step , text prompt , task-specific condition , diffusion algorithm , and noisy image , the objective is as follows.
This objective can also be used directly for fine-tuning. During training, text prompts are randomly replaced with empty strings to improve recognition of the semantic content in the input condition maps.
4. Experiment
See the paper for further experimental results.

4.1. Experiment Setting
All experiments use CFG = 9.0, a DDIM sampler, and 20 steps. They test three prompt types: no prompt, a default prompt ('a professional, detailed, high-quality image'), and an automatic prompt.
5. Limitation

Reference
https://arxiv.org/abs/2302.05543