On Distillation of Guided Diffusion Models
On Distillation of Guided Diffusion Models
0. Abstract
Classifier-free guided diffusion has recently performed well in high-resolution image generation, but inference is slow because both conditional and unconditional models must be evaluated.
The paper proposes distilling classifier-free guided diffusion. First, it trains a single model to produce the combined output of conditional and unconditional models, then distills that model further. Sampling becomes roughly 256 times faster while maintaining FID and IS scores.
Put simply, is slow because both conditional model and unconditional model require evaluation. Distillation speeds it up by combining them, then repeatedly halving the sampling steps—from 1,024 to 4, a 256-fold reduction!
1. Introduction
The two-stage approach first trains one student to match the combined teacher output. It then distills that student into a model requiring fewer sampling steps. A single distilled model supports a broad range of guidance strengths.
Experiments show teacher-like visual quality in four steps and comparable FID/IS scores across guidance strengths in eight to sixteen steps; see Figure 1.

2. Background
Distillation of deep learning models
Suppose we build an AI prediction service using deep learning. Research may produce a complex model trained on extensive data for maximum accuracy. But when deploying it to users, that complexity might be unsuitable.
Which of these two models would be more appropriate?
- Complex model T: 99% prediction accuracy, taking three hours.
- Simple model S: 90% prediction accuracy, taking three minutes.
It depends on the service, but the simpler model S seems more suitable for deployment.
Could we make good use of both? This motivates knowledge distillation: transferring the complex model's learned generalization ability to the simpler model.
The original paper used the terms cumbersome model and simple model. Later research generally calls them teacher and student, reflecting how one learns first and passes its knowledge to the other.
What is classifier-free guidance?
Classifier Guidance
Before diffusion became prominent, GANs improved sample quality at the expense of diversity through truncation or low-temperature sampling. These techniques did not work well for diffusion, initially suggesting worse performance than GANs. Diffusion Models Beat GANs on Image Synthesis changed this by introducing classifier guidance.
Classifier guidance trades sample diversity against fidelity after training a conditional diffusion model. It adds the gradient of an auxiliary classifier's log-likelihood to the sampling score estimate.
In the equation, denotes noisy data.
The classifier term for a particular class is weighted by adjustable w, guiding sampling toward higher-quality images. Its major drawback is requiring class labels and a separately trained classifier.
CFG: Classifier-Free Guidance
Classifier-Free Diffusion Guidance argues that classifier guidance needs noise-level-aware classifiers and can artificially improve classifier-based metrics such as IS and FID, similarly to an adversarial attack.
The researchers propose classifier-free guidance, a simple effective method operating through the diffusion model itself. A one-line training-code change can implement it. The guidance weight w is adjustable.
Here, is conditioned on text , while is unconditional. Rearranging with respect to gives:
The rearranged equation offers an intuitive view of guidance: it increases the difference between the sample's conditional and unconditional likelihoods.
The role of CFG
CFG =
In my original Stable Diffusion experiments, I interpreted CFG as affecting diversity. I usually used 11, observing more prompt-bound images below that and more varied, creative scenes above it.
CFG directly controls the trade-off between sample quality and diversity.
Source: https://pitas.tistory.com/15
3. Distilling a classifier-free guided diffusion model
Given trained guided teacher , the approach has two stages.
Step One,
First, define continuous-time student with trainable parameters to match the teacher at every timestep . Optimize it across guidance-strength range with the objective below.

Here, . To incorporate guidance strength , an -conditioned model takes as input. Initialization matters: the student starts with the conditional teacher's parameters, except for new -conditioning parameters. Details follow.
The model architecture we use is a U-Net model similar to the ones used in Classifier-free diffusion guidance. We use the same number of channels and attention as used in Classifier-free diffusion guidance for both ImageNet 64x64 and CIFAR-10.

Step Two,
Second, distill into fewer-step student , halving sampling steps at each iteration. Given and , with denoting the sampling step, train the student to match two teacher DDIM steps with one step: from to , then to .
After distilling a -step teacher into a -step student, promote the -step student to teacher and repeat, producing a -step student. Each student initializes with its teacher's parameters. Details follow.
We first use the student model from Step-one as the teacher model. We start from 1024 DDIM sampling steps and progressively distill the student model from Step-one to a one step model. We train the student model for 50,000 parameter updates, except for sampling step equals to one or two where we train the model for 100,000 parameter updates, before the number of sampling step is halved and the student model becomes the new teacher model.
I understand how the first iteration reduces two CFG evaluations to one, but not yet why subsequent iterations keep halving the work. Is this enabled mathematically by the objective? If each student replaces two teacher steps with one, perhaps the same equivalence applies repeatedly?

N-step deterministic and stochastic sampling
Once is trained, DDIM sampling is possible for . Note that sampling from distilled model is deterministic for a given initialization .
N-step stochastic sampling is also possible, but evaluates slightly different timesteps and requires a small modification to training compared with the deterministic sampler.
4. Experiments
The approach achieves competitive FID and IS scores in four steps.



Reference
https://arxiv.org/abs/2210.03142
https://baeseongsu.github.io/posts/knowledge-distillation/