DDPM: Diffusion Model Background
The introduction and background of DDPM did not explain diffusion basics sufficiently for me, so I used several blogs to write this separate overview.
0. Abstract1. Introduction2. Background(1) Forward Diffusion Process (q)(2) Reverse Diffusion Process (p)(3) Objective Function (Loss)References
0. Abstract
This paper proposes diffusion probabilistic models, latent-variable models inspired by nonequilibrium thermodynamics that perform high-quality image synthesis. It uses a connection between diffusion and denoising score matching with Langevin dynamics.
1. Introduction
- The paper presents advances in diffusion probabilistic models.
- A diffusion model gradually adds noise to data or reconstructs data from noise to generate data.
- The figure below summarizes this.
is real data, is final noise, and intermediate denotes latent variables containing noisy data.
- The goal is to find . Here, is the reverse process and consists of latents with the same dimensionality as .

- First, the forward process gradually adds noise, moving right to left in the figure. Learning a reverse process that estimates the inverse of this forward process teaches reconstruction of from .
- Using this reverse process, the model can generate images, text, graphs, and other desired data from random noise.
2. Background
(1) Forward Diffusion Process (q)

- The forward diffusion process, q, is a Markov chain that adds noise from the initial data to final noise . This runs opposite to sampling. We need its distribution because reverse-process learning uses information from the forward process.

- Noise is scaled according to the variance schedule before being added.
- Each step samples a Gaussian through reparameterization. Scaling by , rather than merely adding noise, prevents variance from diverging.
- Keeping variance near unit scale maintains it at a controlled level throughout the forward and reverse processes.
- The schedule could be learnable, but experiments found little difference from fixed values, so it is kept constant.
- Its value is small when data still resembles an image and increases as it approaches a Gaussian, rising linearly from 10^-4 to 0.02.
- We can produce from by taking t sequential steps, or do it directly.
- Recursively rearranging the equation gives the distribution of for given data as follows.
- Training one step at a time uses too much memory and computation. Directly producing the noisy sample lets us calculate the loss and take an expectation over t. Since training uses stochastic gradients, this is sufficient.

(2) Reverse Diffusion Process (p)
- The reverse diffusion process p reconstructs from noise.
- Modeling this process is essential because it ultimately generates data from random noise, but determining the actual process is difficult.
We want to know , but it is difficult to obtain.
- We therefore approximate it with , a Markov chain using Gaussian transitions, expressed below.

- At each step, the Gaussian mean and standard deviation are parameters to learn. The starting noise distribution is defined as the simplest standard normal distribution below.
(3) Objective Function (Loss)
- Now that we understand forward and reverse processes, let us examine training to estimate the parameters of .
- Our goal is to find the true data distribution , so we want to maximize likelihood, or minimize negative log-likelihood. This is expressed below.

View proof
: Kullback–Leibler divergence (Kullback–Leibler divergence, KLD)
- Used to measure the difference between two probability distributions.
- The third equality in the training loss holds because both processes are Markov chains: it follows from the Markov property.
Markov property
Given the current state, the probability of the next state is independent of how we reached the current state.
- Finally, rearrange the expression into KL divergences between Gaussian distributions to make computation easier.

- Let us examine each term using the equations introduced earlier.
- LT: The distribution difference between generated by p and generated by given data .
- Lt−1: The difference between the reverse and forward transition distributions p and q. Training makes them as similar as possible.
- L0: The likelihood of estimating data x0 from latent x1, which training maximizes.
and
(1)
(2)
and
(1) is unknown, but is known.
By Bayes' rule, using posterior and prior!
(2)
View proof

- Ultimately, the training loss is easily computed as KL divergence between Gaussian distributions.
References
- Ho, J., Jain, A., & Abbeel, P. (2020). Denoising diffusion probabilistic models. Advances in Neural Information Processing Systems, 33, 6840-6851.