On Distillation of Guided Diffusion Models

On Distillation of Guided Diffusion Models

 
 
 

0. Abstract

Classifier-free guided diffusion has recently performed well in high-resolution image generation, but inference is slow because both conditional and unconditional models must be evaluated.

The paper proposes distilling classifier-free guided diffusion. First, it trains a single model to produce the combined output of conditional and unconditional models, then distills that model further. Sampling becomes roughly 256 times faster while maintaining FID and IS scores.

Put simply,ϵ~t=(1+w)ϵθ(zt,c)−wϵθ(zt)\tilde{\epsilon}_t = (1+w)\epsilon_\theta(z_t,c) - w\epsilon_\theta(z_t) is slow because both conditional model ϵθ(zt,c)\epsilon_\theta(z_t,c) and unconditional model ϵθ(zt)\epsilon_\theta(z_t) require evaluation. Distillation speeds it up by combining them, then repeatedly halving the sampling steps—from 1,024 to 4, a 256-fold reduction!

 

1. Introduction

The two-stage approach first trains one student to match the combined teacher output. It then distills that student into a model requiring fewer sampling steps. A single distilled model supports a broad range of guidance strengths.

Experiments show teacher-like visual quality in four steps and comparable FID/IS scores across guidance strengths in eight to sixteen steps; see Figure 1.

 

2. Background

Distillation of deep learning models

Suppose we build an AI prediction service using deep learning. Research may produce a complex model trained on extensive data for maximum accuracy. But when deploying it to users, that complexity might be unsuitable.

Which of these two models would be more appropriate?

  • Complex model T: 99% prediction accuracy, taking three hours.
  • Simple model S: 90% prediction accuracy, taking three minutes.

It depends on the service, but the simpler model S seems more suitable for deployment.

Could we make good use of both? This motivates knowledge distillation: transferring the complex model's learned generalization ability to the simpler model.

The original paper used the terms cumbersome model and simple model. Later research generally calls them teacher and student, reflecting how one learns first and passes its knowledge to the other.

 
What is classifier-free guidance?

Classifier Guidance

Before diffusion became prominent, GANs improved sample quality at the expense of diversity through truncation or low-temperature sampling. These techniques did not work well for diffusion, initially suggesting worse performance than GANs. Diffusion Models Beat GANs on Image Synthesis changed this by introducing classifier guidance.

Classifier guidance trades sample diversity against fidelity after training a conditional diffusion model. It adds the gradient of an auxiliary classifier's log-likelihood to the sampling score estimate.

In the equation, zλz_\lambda denotes noisy data.

ϵθ(zλ,c)=ϵθ(zλ,c)−wσλ∇zλlog⁡p(c∣zλ)≈−σλ∇zλ[log⁡p(zλ∣c)+wlog⁡pθ(c∣zλ)]\epsilon_\theta(z_\lambda, c) = \epsilon_\theta(z_\lambda,c)-w\sigma_\lambda \nabla_{z_\lambda} \log p(c|z_\lambda) \\ \approx -\sigma_\lambda \nabla_{z_\lambda}[\log p(z_\lambda |c)+w\log p_\theta(c|z_\lambda)]

The classifier term for a particular class is weighted by adjustable w, guiding sampling toward higher-quality images. Its major drawback is requiring class labels and a separately trained classifier.

 

CFG: Classifier-Free Guidance

Classifier-Free Diffusion Guidance argues that classifier guidance needs noise-level-aware classifiers and can artificially improve classifier-based metrics such as IS and FID, similarly to an adversarial attack.

The researchers propose classifier-free guidance, a simple effective method operating through the diffusion model itself. A one-line training-code change can implement it. The guidance weight w is adjustable.

Here, ϵθ(zt,c)\epsilon_\theta(z_t, c) is conditioned on text cc, while ϵθ(zt)\epsilon_\theta(z_t) is unconditional. Rearranging with respect to ww gives:

The rearranged equation offers an intuitive view of guidance: it increases the difference between the sample's conditional and unconditional likelihoods.

 

The role of CFG

CFG = w∈R≥0w \in \mathbb{R}^{\geq 0}

In my original Stable Diffusion experiments, I interpreted CFG as affecting diversity. I usually used 11, observing more prompt-bound images below that and more varied, creative scenes above it.

CFG directly controls the trade-off between sample quality and diversity.

 
 

3. Distilling a classifier-free guided diffusion model

Given trained guided teacher [x^c,θ,x^θ][\hat{x}_{c,\theta}, \hat{x}_\theta], the approach has two stages.

Step One,

First, define continuous-time student x^η1(zt,w)\hat{x}_{\eta_1}(z_t,w) with trainable parameters η1\eta_1 to match the teacher at every timestep t∈[0,1]t \in [0,1]. Optimize it across guidance-strength range [wmin,wmax][w_{min}, w_{max}] with the objective below.

Here, x^θw(zt)=(1+w)x^c,θ(zt)−wx^θ(zt)\hat{x}_\theta^w(z_t) = (1+w)\hat{x}_{c,\theta}(z_t) - w\hat{x}_\theta(z_t). To incorporate guidance strength ww, an ww-conditioned model takes ww as input. Initialization matters: the student starts with the conditional teacher's parameters, except for new ww-conditioning parameters. Details follow.

The model architecture we use is a U-Net model similar to the ones used in Classifier-free diffusion guidance. We use the same number of channels and attention as used in Classifier-free diffusion guidance for both ImageNet 64x64 and CIFAR-10.

 

Step Two,

Second, distill x^η1(zt,w)\hat{x}_{\eta_1}(z_t,w) into fewer-step student x^η2(zt,w)\hat{x}_{\eta_2}(z_t,w), halving sampling steps at each iteration. Given w∼U[wmin,wmax]w \sim U[w_{min}, w_{max}] and t∈[1,…,N]t \in [1,…,N], with NN denoting the sampling step, train the student to match two teacher DDIM steps with one step: from t/Nt/N to t−0.5/Nt-0.5/N, then t−0.5/Nt - 0.5/N to t−1/Nt-1/N.

After distilling a 2N2N-step teacher into a NN-step student, promote the NN-step student to teacher and repeat, producing a N/2N/2-step student. Each student initializes with its teacher's parameters. Details follow.

We first use the student model from Step-one as the teacher model. We start from 1024 DDIM sampling steps and progressively distill the student model from Step-one to a one step model. We train the student model for 50,000 parameter updates, except for sampling step equals to one or two where we train the model for 100,000 parameter updates, before the number of sampling step is halved and the student model becomes the new teacher model.

I understand how the first iteration reduces two CFG evaluations to one, but not yet why subsequent iterations keep halving the work. Is this enabled mathematically by the objective? If each student replaces two teacher steps with one, perhaps the same equivalence applies repeatedly?

 

N-step deterministic and stochastic sampling

Once x^η2\hat{x}_{\eta_2} is trained, DDIM sampling is possible for w∈[wmin,wmax]w \in [w_{min}, w_{max}]. Note that sampling from distilled model x^η2\hat{x}_{\eta_2} is deterministic for a given initialization z1wz_1^w.

N-step stochastic sampling is also possible, but evaluates slightly different timesteps and requires a small modification to training compared with the deterministic sampler.

 

4. Experiments

The approach achieves competitive FID and IS scores in four steps.

 

Reference

https://arxiv.org/abs/2210.03142

https://baeseongsu.github.io/posts/knowledge-distillation/

https://pitas.tistory.com/15

 
 

Read next