DDPM: Diffusion Model Background

The introduction and background of DDPM did not explain diffusion basics sufficiently for me, so I used several blogs to write this separate overview.
 
 

0. Abstract

This paper proposes diffusion probabilistic models, latent-variable models inspired by nonequilibrium thermodynamics that perform high-quality image synthesis. It uses a connection between diffusion and denoising score matching with Langevin dynamics.

 

1. Introduction

  • The paper presents advances in diffusion probabilistic models.
  • A diffusion model gradually adds noise to data or reconstructs data from noise to generate data.
  • The figure below summarizes this. x0x_0 is real data, xTx_T is final noise, and intermediate xtx_t denotes latent variables containing noisy data.
  • The goal is to find pθ(x0)=∫pθ(x0:T)dx1:Tp_{\theta}(x_0) = \int p_{\theta}(x_{0:T}) dx_{1:T}. Here, pθ(x0:T)p_{\theta}(x_{0:T}) is the reverse process and x1,…,xTx_1,…,x_T consists of latents with the same dimensionality as x0∼q(x0)x_0 \sim q(x_0).
  • First, the forward process q(xt∣xt−1)q(x_{t}|x_{t-1}) gradually adds noise, moving right to left in the figure. Learning a reverse process pθ(xt−1∣xt)p_{\theta}(x_{t-1}|x_t) that estimates the inverse of this forward process teaches reconstruction of data(x0)data(x_0) from noise(xT)noise(x_T).
  • Using this reverse process, the model can generate images, text, graphs, and other desired data from random noise.
 

2. Background

(1) Forward Diffusion Process (q)

  • The forward diffusion process, q, is a Markov chain that data(x0)data(x_0)adds noise from the initial data to final noise noise(xT)noise(x_T). This runs opposite to sampling. We need its distribution because reverse-process learning uses information from the forward process.
 
forward process의 수식 표현
The forward-process equation
  • Noise is scaled according to the variance schedule {βt∈(0,1)}t=1T\{\beta_t \in (0,1)\}_{t=1}^{T} before being added.
    • Each step samples a Gaussian through reparameterization. Scaling by 1−βt\sqrt{1-\beta_t}, rather than merely adding noise, prevents variance from diverging.
    • Keeping variance near unit scale maintains it at a controlled level throughout the forward and reverse processes.
      • xt=1−βt xt−1+βt∗ϵ(1−βt)2+βt=1x_t = \sqrt{1-\beta_t}\ x_{t-1} + \beta_t * \epsilon \\ (\sqrt{1-\beta_t})^2 + \beta_t = 1
    • The schedule could be learnable, but experiments found little difference from fixed values, so it is kept constant.
    • Its value is small when data still resembles an image and increases as it approaches a Gaussian, rising linearly from 10^-4 to 0.02.
  • We can produce xtx_t from x0x_0 by taking t sequential steps, or do it directly.
    • Recursively rearranging the equation gives the distribution of xtx_t for given data x0x_0 as follows.
    • at:=1−βt and αtˉ:=∏s=1tαsq(xt∣x0)=N(xt;αtˉx0, (1−αtˉ)I)a_t := 1-\beta_t \ and \ \bar{\alpha_t} := \prod_{s=1}^{t}\alpha_s \\ q(x_t|x_0) = \mathcal{N}(x_t;\sqrt{\bar{\alpha_t}}x_0,\ (1-\bar{\alpha_t})\mathrm{I})
    • Training one step at a time uses too much memory and computation. Directly producing the noisy sample lets us calculate the loss and take an expectation over t. Since training uses stochastic gradients, this is sufficient.
 
reparameterize 증명
Reparameterization proof
 

(2) Reverse Diffusion Process (p)

  • The reverse diffusion process p noise(xT)noise(x_T) reconstructs data(x0)data(x_0) from noise.
  • Modeling this process is essential because it ultimately generates data from random noise, but determining the actual process is difficult.
    • We want to know q(xt−1∣xt)q(x_{t-1}|x_t), but it is difficult to obtain.

  • We therefore approximate it with pθpθ, a Markov chain using Gaussian transitions, expressed below.
reverse process의 수식 표현
The reverse-process equation
  • At each step, the Gaussian mean μθ\mu_θ and standard deviation ∑θ\sum_θ are parameters to learn. The starting noise distribution is defined as the simplest standard normal distribution below.
p(xT)=N(xT;0,I)p(x_T) = \mathcal{N}(x_T;0,\mathrm{I})
 

(3) Objective Function (Loss)

  • Now that we understand forward and reverse processes, let us examine training to estimate the parameters of pθp_\theta.
  • Our goal is to find the true data distribution pθ(x0)p_\theta(x_0), so we want to maximize likelihood, or minimize negative log-likelihood. This is expressed below.
Diffusion Model의 Loss Function
The diffusion-model loss function
View proof
−log⁡pθ(x0)≤−log⁡pθ(x0)+DKL(q(x1:T∣x0)∣∣pθ(x1:T∣x0))=−log⁡pθ(x0)+Ex1:T∼q(x1:T∣x0)[log⁡q(x1:T∣x0)pθ(x0:T)/pθ(x0)]=−log⁡pθ(x0)+Ex1:T∼q(x1:T∣x0)[log⁡q(x1:T∣x0)pθ(x0:T)+log⁡pθ(x0)]=Eq[log⁡q(x1:T∣x0)pθ(x0:T)]=Eq[−log⁡pθ(x0:T)q(x1:T∣x0)]-\log {p_\theta}(x_0) \leq -\log {p_\theta}(x_0) + D_{KL}(q(x_{1:T}|x_0)||p_{\theta}(x_{1:T}|x_0)) \\ = -\log p_\theta(x_0) + \mathbb{E}_{x_{1:T}\sim q(x_{1:T}|x_0)}[\log {q(x_{1:T}|x_0) \over p_{\theta}(x_{0:T})/p_\theta(x_0)}] \\ = -\log p_\theta(x_0) + \mathbb{E}_{x_{1:T}\sim q(x_{1:T}|x_0)}[\log {q(x_{1:T}|x_0) \over p_{\theta}(x_{0:T})}+\log p_\theta(x_0)] \\ = \mathbb{E}_q[\log{q(x_{1:T}|x_0) \over p_\theta(x_{0:T})}] \\ = \mathbb{E}_q[-\log{ p_\theta(x_{0:T})\over q(x_{1:T}|x_0)}]
DKLD_{KL}: Kullback–Leibler divergence (Kullback–Leibler divergence, KLD)
  • Used to measure the difference between two probability distributions.
  • DKL(P∣∣Q)=∑iP(i)log⁡P(i)Q(i)D_{KL}(P||Q) = \sum_iP(i) \log {P(i) \over Q(i)}
 
  • The third equality in the training loss holds because both processes are Markov chains: it follows from the Markov property.
    • Markov property

      Given the current state, the probability of the next state is independent of how we reached the current state.

  • Finally, rearrange the expression into KL divergences between Gaussian distributions to make computation easier.
Diffusion Model의 최종적인 Loss Function
The final diffusion-model loss
  • Let us examine each term using the equations introduced earlier.
    • LT: The distribution difference between noise(xT)noise(x_T) generated by p and noise(xT)noise(x_T) generated by qq given data x0x_0.
      • q(xT∣x0)q(x_T|x_0) and p(xT)p(x_T)

        (1) q(xt∣x0)=N(xt;αtˉx0, (1−αtˉ)I)q(x_t|x_0) = \mathcal{N}(x_t;\sqrt{\bar{\alpha_t}}x_0,\ (1-\bar{\alpha_t})\mathrm{I})

        (2) p(xT)=N(xT;0,I)p(x_T) = \mathcal{N} (x_T;0,\mathrm{I})

    • Lt−1: The difference between the reverse and forward transition distributions p and q. Training makes them as similar as possible.
      • q(xt−1∣xt,x0)q(x_{t-1}|x_t, x_0) and pθ(xt−1∣xt)p_{\theta}(x_{t-1}|x_t)
        (1) q(xt−1∣xt)q(x_{t-1}|x_t) is unknown, but q(xt−1∣xt,x0)q(x_{t-1}|x_t, x_0) is known.
        By Bayes' rule, using posterior and prior! P(xt−1∣xt)=P(xt∣xt−1)P(xt−1)P(xt)P(x_{t-1}|x_t) = {P(x_t|x_{t-1}) P(x_{t-1}) \over P(x_t)}

        q(xt−1∣xt,x0)=q(xt∣xt−1,x0)q(xt−1∣x0)q(xt∣x0)q(x_{t-1}|x_t, x_0) = q(x_t|x_{t-1}, x_0){q(x_{t-1}|x_0)\over q(x_t|x_0)}

        ∴reverse conditional probability, q(xt−1∣xt,x0)=N(xt−1;μ~(xt,x0),β~tI)μ~t(xt,x0)=αˉt−11−αtˉβtx0 +αt(1−αˉt−1)1−αˉtxt, βt~=(1−αˉt−1)1−αˉt\therefore reverse\ conditional \ probability,\\ \ q(x_{t-1}|x_t, x_0) = N(x_{t-1}; \tilde{\mu}(x_t,x_0), \tilde{\beta}t\mathrm{I}) \\ \tilde{\mu}t(x_t,x_0) = \sqrt{\bar{\alpha}{t-1}\over1-\bar{\alpha_t}}\beta_tx_0 \ + {\sqrt{\alpha_t}(1-\bar{\alpha}{t-1})\over 1-\bar{\alpha}t}x_t, \ \tilde{\beta_t} = {(1-\bar{\alpha}{t-1})\over 1-\bar{\alpha}_t}

        (2) pθ(xt−1∣xt):=N(xt−1;μθ(xt,t), ∑θ(xt,t))p_{\theta}(x_{t-1}|x_t):=\mathcal{N}(x_{t-1};\mu_{\theta}(x_t,t),\ \sum_\theta (x_t,t))

    • L0: The likelihood of estimating data x0 from latent x1, which training maximizes.
      •  
View proof
최종 Loss Function 증명
Proof of the final loss function
 
  • Ultimately, the training loss is easily computed as KL divergence between Gaussian distributions.
 

References

  1. Ho, J., Jain, A., & Abbeel, P. (2020). Denoising diffusion probabilistic models.  Advances in Neural Information Processing Systems, 33, 6840-6851.
  1. https://lilianweng.github.io/posts/2021-07-11-diffusion-models/
  1. https://happy-jihye.github.io/diffusion/diffusion-1/
  1. https://process-mining.tistory.com/182
 
 

Read next