DDPM: Denoising Diffusion Probabilistic Models
DDPM: Denoising Diffusion Probabilistic Models
1. Introduction
In short, DDPM gradually adds small Gaussian noise, with fixed time-dependent parameters, to an image until only normally distributed noise remains. The problem is how to reconstruct an image from that noise. Recovering an image from supplied noise becomes image generation.
The paper describes diffusion probabilistic models: parameterized Markov chains trained through variational inference to generate images after a finite number of steps. A Markov chain has the form , where the previous sample influences the current one. In one direction, small Gaussian noise is progressively added until the image is destroyed.
2. Background

2-1. Forward Diffusion Process ()
Let the original image be and the noising process . Applying noise to to obtain is written , or generally at time t as . This is the forward, or diffusion, process. It ends at , pure noise following normal distribution . For equations, see the previous post.
2-2. Reverse Diffusion Process ()
The reverse process gradually removes noise, opposite to . Reversing time in gives . Note that all have the same resolution. For equations, see the previous post.
2-3. Objective Function ()
The objective is to gradually remove noise from a given noisy sample, using .
Given if we can predict this, we can also predict the following.
This operates through Gaussians parameterized by mean and variance. As a generative model, it maximizes generated-image log-likelihood, a measure of fit to the data distribution. The paper expresses this as minimizing negative log-likelihood.
Loss Function

The equation gives the variational bound on the right-hand side. Expanding it yields:

This is the resulting expression.
Notably, it directly compares the forward-process posterior and reverse process through KL divergence. Since the forward process is known and its posterior is closely related to the reverse process, the comparison is tractable.
A brief look at what posterior probability means
Link: https://hwiyong.tistory.com/27
1. Posterior
Posterior means probability after observation. Its counterpart is the prior; likelihood is another related quantity.

Suppose we observe a silhouette on a curtain. Consider two probabilities:
P(silhouette | Cheolsu) and P(Cheolsu | silhouette).
What we actually want is P(Cheolsu | silhouette).
This is the posterior, but we cannot usually calculate it directly. Variational inference and related methods
choose to approximate it rather than directly calculate the probability.
We use P(silhouette | Cheolsu), the likelihood, in that approximation.
- Inferring whether a silhouette belongs to Cheolsu may feel intuitive, but the reverse conditional can be less intuitive.
- P(silhouette | Cheolsu) considers the possible silhouettes Cheolsu could produce—animal-like, human-like, airplane-like, and so on.
- It gives the probability of observing the particular silhouette on the curtain among those possibilities.
Bayes' rule relates these quantities to the posterior.
The expression looks like conditional probability because it is the same concept.
We often rearrange it once more when approximating the quantity we want.
The denominator P(silhouette) is constant with respect to the alternatives in the numerator.
Ultimately,
we approximate in the form below. Using more formal mathematical notation,
we obtain this expression!
Known probabilities such as P(Cheolsu) and P(Yeonghui) are called priors.
Bayesian inference therefore obtains the posterior using likelihood and prior.
☞ Likelihood: p(z | x), the probability of observed data under a model.
☞ Prior probability: p(x), our probability for a system or model before observation. Population proportions such as p(man) and p(woman) are examples.
☞ Posterior probability: p(x | z), the probability of a model after observing an event.
Now consider the variational bound above.
3. Diffusion models and denoising autoencoders
3-1. Objective Function ()

- The most important term is . Starting from and expanding conditionally gives the tractable Gaussian forward posterior . Its KL divergence allows us to train the desired .
3-2. Forward Diffusion Process ()
- The objective's first term
- Forward variances are fixed constants rather than learnable parameters, so can be ignored.
- Why?
- The forward process makes Gaussian, so tractable distribution is very close to prior . Since forward variance is fixed, this approximate posterior contains no learnable parameters.
- Thus, this loss term is a constant near zero and is ignored during training.
3-3. Reverse Diffusion Process ()
- The objective's second term
- First determine 's distribution, then obtain and to determine . We will examine them in order.
- (1)
As in the earlier diffusion overview, it is defined in the following form.
- (2)
As in the earlier diffusion overview, it is defined in the following form.
- (3)
The standard deviation of is a constant matrix , so it need not be learned. Although is more properly expressed as , using is also acceptable.
- (4)
Finally, the mean of is defined as follows.


Proof

- (5) Sampling
Once we obtain , we can sample from with the algorithm below.

- (6)
Combining , , and to calculate , the KL divergence, gives as follows.
The expectation changes from q to and . This is natural because the forward Gaussian at any time can be reparameterized given the time and .



- (7)
Finally, rewrite the loss in terms of epsilon. This simplified objective improves training.
The simplified objective is effective because it trains the network at large t as well as very small t.


- (8) Training
DDPM training can therefore be summarized in the following short algorithm.

- (9)
The final loss component is a simple KL divergence between two normal distributions and has the form below.
4. Experiments
The central contribution is a convenient reparameterization of reverse-process . Defining forward and reverse processes this way produces a generative model that synthesizes high-quality images.


References
- Ho, J., Jain, A., & Abbeel, P. (2020). Denoising diffusion probabilistic models. Advances in Neural Information Processing Systems, 33, 6840-6851.