ControlNet: Adding Conditional Control to Text-to-Image Diffusion Models

ControlNet: Adding Conditional Control to Text-to-Image Diffusion Models

 
 

0. Abstract

This paper introduces ControlNet, which controls a pretrained diffusion model such as Stable Diffusion using additional inputs: edge maps, segmentation maps, keypoints, and more. It learns task-specific conditions and trains robustly even on datasets smaller than 50,000 examples.

1. Introduction

After writing text to generate an image, we might ask:

  1. Does text-prompt-based control adequately meet our need to generate the image we imagined?
  1. What framework should we build to accommodate a wide range of task conditions and user controls, including inputs other than text?
  1. Can we preserve a large model's capabilities learned from many images when adapting it to a specific task?
 

To answer these questions, the authors surveyed image-processing applications and made three observations, all pointing toward the need for less data and computation.

Three observations
  1. Task-specific datasets are not always as large as general text-to-image datasets.
      • Most task-specific datasets contain fewer than 100,000 examples—5×1045\times10^4 smaller than LAION-5B.
  1. Large computational resources are not always available for data-driven image-processing solutions.
      • Fast training methods are therefore important when optimizing large models, ideally allowing training on personal devices.
  1. Image-processing problems take many forms in their definitions, user controls, and image annotations.
      • Tasks such as depth-to-image and pose-to-human require object- or scene-level interpretation of raw inputs. Building a procedural method for each one—constraining denoising, modifying multi-head attention activations, and so on—is difficult. End-to-end learning is essential for learned solutions across many tasks.
 
 

ControlNet is an end-to-end neural architecture that controls large image diffusion models such as Stable Diffusion with specific input conditions. It duplicates model weights into a trainable copy and a locked copy. The trainable copy learns conditional control from task-specific data, while the locked copy preserves the original capabilities. Special layers called zero convolutions connect the two blocks. Preserving production-ready weights makes training robust across dataset sizes. Because zero convolutions introduce no new noise, training is also as fast as fine-tuning the diffusion model.

 

Effects of a 1×1 convolution

  • Reducing trainable parameters by changing channel counts
  • Adding nonlinearity
  • Computing object and positional information independently
 

3. Method

3.1. ControlNet

ControlNet manipulates the input conditions of neural network blocks to control the overall network's behavior. Such blocks include ResNet blocks, conv–BN–ReLU blocks, multi-head attention blocks, and transformer blocks.

For 2D features, given a feature map x∈Rh×w×cx \in \mathbb{R}^{h\times w \times c}, a neural network block F(⋅ ;Θ)\mathcal{F}(\cdot \ ; \mathrm{\Theta}), parameterized by Θ\Theta, transforms xx into another feature map yy.

All parameters in Θ\Theta are locked and cloned into a trainable copy Θc\Theta_c. The copied Θc\Theta_c is trained using an external condition vector cc. The paper calls these the locked copy and trainable copy. Rather than directly updating the original weights, this preserves the quality learned from hundreds of millions of images and avoids overfitting on small datasets.

The blocks are connected by a special convolution layer Z(⋅;⋅)\mathcal{Z}(\cdot;\cdot), called a zero convolution: a 1×1 convolution whose weights and biases are initialized to zero. Z\mathcal{Z} uses two parameter instances {Θz1,Θz2}\{\Theta_{z_1}, \Theta_{z_2}\}, producing the ControlNet structure below.

A trained ControlNet recognizes the semantic content of its input conditions. In other words, it reflects the meaning of depth, normals, edges, and other conditions in generation.

yc=F(x;Θ)+Z(F(x+Z(c;Θz1);Θc);Θz2)y_c=\mathcal{F}(x;\Theta) + \mathcal{Z}(\mathcal{F}(x+\mathcal{Z}(c;\Theta_{z_1});\Theta_c); \Theta_{z_2})

Because zero-convolution weights and biases start at zero, the first training step gives the following equations.

We can interpret them as follows.

yc=yy_c=y

These equations show that adding ControlNet does not affect the original network block before optimization. Its capabilities, functionality, and output quality remain intact. Subsequent optimization is as fast as fine-tuning, and much faster than training these layers from scratch.

Let us briefly examine gradients in a zero convolution. For a 1×1 convolution with weight WW and bias BB, spatial position h×wh\times w denoted by pp, and channel index ii, the forward pass for an input map I∈Rh×w×cI \in \mathbb{R}^{h\times w \times c} is as follows.

Before optimization, zero convolution has W=0W=0 and B=0B=0. Wherever Ip,iI_{p,i} is nonzero, its gradients are:

Although zero convolution can make the gradient for the feature term II zero, the gradients for weight WW and bias BB are unaffected. At the first gradient-descent step, the weight WW becomes a nonzero matrix as long as the feature II is nonzero. Simply put,WW even if BB are initialized to zero, II only needs to be nonzero for learning to occur.

Thus, zero convolution becomes a distinctive connection layer that evolves from zero into optimized parameters.

 

3.2. ControlNet in Image Diffusion Model

Stable Diffusion is used as the example.

To stabilize training, Stable Diffusion preprocesses its 512×512 images into 64×64 latent images, similarly to VQ-GAN. ControlNet must likewise encode image-based conditions into a 64×64 feature space to match convolution sizes. The paper uses a small network E(⋅)\mathcal{E}(\cdot), with 4×4 kernels and 2×2 strides, to encode an image-space condition cic_i into a feature map cfc_f.

The encoder transforms a 512×512 condition into a 64×64 feature map.

 

3.3. Training

Given a time step tt, text prompt ctc_t, task-specific condition cfc_f, diffusion algorithm ϵθ\epsilon_\theta, and noisy image ztz_t, the objective is as follows.

L=Ez0,t,ct,cf,ϵ∼N(0,1)[∣∣ϵ−ϵθ(zt,t,ct,cf)∣∣22]\mathcal{L} = \mathbb{E}_{z_0,t,c_t,c_f,\epsilon \sim \mathcal{N}(0,1)}\bigg[|| \epsilon-\epsilon_{\theta}(z_t,t,c_t,c_f) ||_2^2\bigg]

This objective can also be used directly for fine-tuning. During training, text prompts ctc_t are randomly replaced with empty strings to improve recognition of the semantic content in the input condition maps.

 

4. Experiment

See the paper for further experimental results.

4.1. Experiment Setting

All experiments use CFG = 9.0, a DDIM sampler, and 20 steps. They test three prompt types: no prompt, a default prompt ('a professional, detailed, high-quality image'), and an automatic prompt.

 

5. Limitation

 

Reference

https://arxiv.org/abs/2302.05543

 

Read next