DDIM: Denoising Diffusion Implicit Models

DDPM teaches a generative Markov chain that removes noise gradually. Sequential sampling needs many steps and is much slower than single-pass GANs. DDIM generalizes that process into a non-Markovian one.
 

DDIM: Denoising Diffusion ImplicitModels

 
 

0. Abstract

DDPM generates high-quality images without adversarial training, but requires a long Markov chain and many sampling steps.

Denoising diffusion implicit models (DDIM) accelerate sampling by generalizing DDPM to non-Markovian diffusion. This enables a more deterministic generative process and high-quality samples ten to fifty times faster.

 

1. Introduction

Deep generative models have demonstrated high-quality generation across image domains. GANs outperform likelihood-based VAEs and flows in quality, but stable training restricts architectures and optimization, and they fail to cover all distribution modes.

DDPM and noise-conditional score networks produce GAN-quality samples without adversarial training by denoising through a generative Markov chain. Many iterations make them much slower than single-pass GANs.

To narrow this efficiency gap, DDIM uses the same training objective as DDPM but generalizes the forward process to non-Markovian chains. This can improve sampling efficiency ten to one hundred times with only a small quality loss.

DDIM also offers stronger consistency: samples starting from the same latent have similar high-level features. This enables semantically meaningful interpolation.

 

2. Background

If all conditionals are Gaussian with learnable means and fixed variances, DDPM's objective simplifies as follows.

DDPM의 obejective function
DDPM's objective function

Let ϵθ:={ϵθ(t)}t=1T\epsilon_\theta := \{ \epsilon_\theta^{(t)} \}_{t=1}^T be a collection of TT functions. Each timestep-indexed ϵθ(t):X→X\epsilon_\theta^{(t)} : \mathcal{X} \rightarrow \mathcal{X} depends on trainable parameters θ(t)\theta^{(t)} and positive-coefficient vector γ:=[γ1,...,γT]\gamma := [\gamma_1,...,\gamma_T], whose coefficients depend on α1:T\alpha_{1:T}. DDPM optimizes objective γ=1\gamma = 1 for generation performance. To generate x0x_0, sample xTx_T from prior pθ(xT)p_\theta(x_T), then repeatedly sample xt−1x_{t-1}.

Forward-process length T is an important hyperparameter. Larger T makes reverse transitions closer to Gaussian and improves their approximation. DDPM therefore uses large T, such as 1,000. But these T steps are sequential, not parallel,x0x_0 making sampling slower than other generative models.

 

3. Variational Inference for Non-Markovian Forward Process

Reconsidering inference, the authors find that objective LγL_\gamma depends only on marginal distributions q(xt∣x0)q(x_t|x_0), not directly on joint distribution q(x1:T∣x0)q(x_{1:T}|x_0). Many joints share those marginals, allowing a new non-Markovian process, shown on Figure 1's right. They establish this for Gaussian cases.

Definitions of joint and marginal distributions
 

3.1. Non-Markovian Forward Process

DDPM과 DDIM의 분포
DDPM and DDIM distributions

To satisfy these equations, the forward inference mean must be:

qσ(xt−1∣xt,x0)=N(αt−1x0+1−αt−1−σt2xt−αtx01−αt,σt2I)q_\sigma(x_{t-1}|x_t, x_0) = \mathcal{N}(\sqrt{\alpha_{t-1}}x_0 + \sqrt{1-\alpha_{t-1}-\sigma_t^2} \cfrac{x_t - \sqrt{\alpha_t}x_0}{\sqrt{1-\alpha_t}}, \sigma_t^2I)

The mean ensures qσ(xt∣x0)=N(αtx0,(1−αt)I)q_\sigma(x_t|x_0) = \mathcal{N}(\sqrt{\alpha_t}x_0, (1-\alpha_t)I) at every t, so the joint distribution has the desired marginals. Bayes' rule gives the following forward process.

qσ(xt∣xt−1,x0)=qσ(xt−1∣xt,x0)qσ(xt∣x0)qσ(xt−1∣x0)q_\sigma(x_t |x_{t-1}, x_0) = \cfrac{q_\sigma(x_{t-1}|x_t,x_0)q_\sigma(x_t|x_0)}{q_\sigma(x_{t-1}|x_0)}

Unlike DDPM, DDIM's forward process is no longer Markovian, because xtx_t depends on both xt−1x_{t-1} and x0x_0. σ\sigma controls forward-process stochasticity. As it approaches zero,xt−1x_{t-1} becomes fixed and deterministic.

 

3.2. Generative Process and Unified Variational Inference Objective

Next define learnable pθ(x0:T)p_\theta(x_{0:T}). pθ(t)(xt−1∣xt)p_\theta^{(t)}(x_{t-1}|x_t) uses known reverse conditional qσ(xt−1∣xt,x0)q_\sigma(x_{t-1} |x_{t}, x_0) to sample xt−1x_{t-1} from noisy xtx_t using x0x_0. Intuitively, for observation xtx_t, first predict corresponding x0x_0, then sample xt−1x_{t-1} through qσ(xt−1∣xt,x0)q_\sigma(x_{t-1} |x_{t}, x_0).

Given x0∼q(x0)x_0 \sim q(x_0) and ϵt∼N(0,I)\epsilon_t \sim \mathcal{N}(0,I), xt=αtx0+1−αtϵx_t = \sqrt{\alpha_t}x_0 + \sqrt{1-\alpha_t}\epsilon Equation 4 gives xtx_t. Model ϵθ(t)(xt)\epsilon_\theta^{(t)}(x_t) predicts ϵt\epsilon_t from xtx_t without knowing x0x_0. Rearranging for x0x_0 yields x0x_0's prediction for xtx_t, the denoised observation. Fixed prior pθ(xT)=N(0,I)p_\theta(x_T) = \mathcal{N}(0,I) then defines the generative process.

 

(1) denoised observation

fθ(t)(xt)=1αt(xt−1−αtϵθ(t)(xt)) f_\theta^{(t)}(x_t) = \cfrac{1}{\sqrt{\alpha}_t}(x_t - \sqrt{1-\alpha_t}\epsilon_\theta^{(t)}(x_t))
  • Expressing Equation 4 in terms of x0x_0
  • Predicting x0x_0 with xtx_t
 

(2) generative process

⁍⁍
  • In qσ(xt−1∣xt, fθ(t)(xt)q_\sigma(x_{t-1}|x_t, \ f_\theta^{(t)}(x_t), denoised observation ff replaces x0x_0.
  • Gaussian noise σ12I\sigma_1^2I is added when t=1t=1 to ensure a generative process.
 

(3) variational inference objective

  • Different choices of σ\sigma yield different models.
    • Setting σ=0\sigma = 0 gives the DDIM objective.
    • Setting σ≠0\sigma ≠ 0 gives the DDPM objective.
 

(4) Theorem 1.

For all σ>0 there exists γ∈R>0T and C∈R, such that Jσ=Lγ+CFor \ all \ \sigma > 0 \, there \ exists \ \gamma \in \mathbb{R}_{>0}^T \ and \ C \in \mathbb{R}, \\ \ such \ that \ J_\sigma = L_\gamma + C
  • If γ=1\gamma = 1, the objective equals DDPM's variational lower bound.
    • L1=JσL_1 = J_\sigma
  • Theorem 1 shows that JσJ_\sigma is a form of LγL_\gamma; JσJ_\sigma has its optimum when γ=1\gamma=1 (L1L_1).
DDPM의 obejective function
DDPM's objective function
 

4. Sampling from Generalized Generative Process

Since γ=1\gamma=1 gives the optimum, use L1L_1 as the objective.

Choosing σ\sigma yields Markovian or non-Markovian processes. Regardless of σ\sigma, the parameters to learn are θ\theta.

  • Thus, pretrained DDPM parametersθ\theta can also be used in DDIM's generative process.
  • Rather than introducing a new training method, DDIM generalizes diffusion to a non-Markovian chain and introduces a faster sampling method.
  • A common approach is training with DDPM and sampling with DDIM, combining a strong model with fast generation.
 

4.1. Denoising Diffusion Implicit Model

In pθ(x1:T)p_\theta(x_{1:T}), the equation below generates sample xt−1x_{t-1} from xtx_t. ϵt∼N(0,I)\epsilon_t \sim \mathcal{N}(0,I) is standard Gaussian noise independent of xtx_t, and α0:=1\alpha_0 := 1 defines the setting. Different σ\sigma yield different generative processes, but shared ϵθ\epsilon_\theta means no retraining.

 
(1) if, σt=1−αt−11−αt1−αtαt−1\sigma_t = \sqrt{\cfrac{1-\alpha_{t-1}}{1-\alpha_t}}\sqrt{\cfrac{1-\alpha_t}{\alpha_{t-1}}} for all tt
  • forward process = “Markovian” = DDPM
 
(2) if, σt=0\sigma_t = 0 for all tt
  • The forward process becomes deterministic with respect to xt−1x_{t-1} and x0x_0, except t=1t=1.
  • σtϵt\sigma_t\epsilon_t becomes zero, making the model an implicit probabilistic model.
  • The sample-generation process from xTx_T to x0x_0 becomes fixed, allowing faster sampling.
  • This is DDIM: An implicit probabilistic model trained with the DDPM objective.
    • The forward process is no longer diffusion in this case.

4.2. Accelerated Generation Processes

Previously, generation approximated the reverse of a TT-step forward process, requiring TT sampling steps. But denoising objective L1L_1, with qσ(xt∣x0)q_\sigma(x_t|x_0) fixed, no longer depends on that forward trajectory. We can consider fewer than TT steps, accelerating generation without training another model.

Suppose the forward process is defined not over all latents x1:Tx_{1:T}, but a subset{xτ1,...,xτS}\{x_{\tau_1},..., x_{\tau_{S}}\}. τ\tau is an increasing subsequence of [1,…,T][1,…,T] of length S. Define its transitions so q(xτi∣x0)q(x_{\tau_i}|x_0) matches the marginals.

Generation samples in reverseτ\tau. This is the sampling trajectory. If the trajectory is shorter than TT, computational efficiency improves substantially.

Small changes to Equation 12 yield a faster process applicable to DDPM and DDIM. Subsampled trajectories generate images faster, but if a model trains only at selected forward steps, generation can use only those steps; broader timestep training is needed.

  • Training at many steps as in DDPM is more effective than training only selected steps in this interpretation of DDIM.
  • Hence DDPM training with DDIM sampling is common.
 

4.3. Relevance to Neural ODEs

What is an ordinary differential equation?

Rewriting DDIM iterations as Equation 12 makes their similarity to Euler integration for ordinary differential equations clearer.

To derive the ODE, reparameterize (1−α/α\sqrt{1-\alpha}/\sqrt{\alpha}) with σ\sigma and (x/αx/\sqrt{\alpha}) with xˉ\bar{x}. When σ(0)=0\sigma(0) = 0, Equation 13 becomes Euler's method for the ODE below.

With sufficiently fine discretization, reversing Equation 14's ODE enables encoding, reversing generation: x0x_0 → xTx_T.

 

5. Experiments

DDIM needs far fewer iterations than DDPM, giving ten- to one-hundred-fold speedups. With initial latent xTx_T fixed, different trajectories preserve high-level features, and latent interpolation becomes possible.

DDIM can encode samples and reconstruct images from latents, unlike DDPM's stochastic process. Each dataset uses the same trained model with T=1000T=1000 and objective L1L_1. Only how samples are produced changes.

Experiments vary subsequences τ\tau of [1,...,T][ 1, ... , T ] and variance hyperparameter σ\sigma. η∈R≥0\eta \in \mathbb{R}_{\ge0} is directly controllable.

  • If η=1\eta=1: DDPM.
  • If η=0\eta=0: DDIM.
 

5.1. Sample Quality and Efficiency

  • Figure 3 reports FID for models trained on CIFAR-10 and CelebA.
  • Quality improves as dim(τ)dim(\tau) increases, trading off against computation. For the same dim(τ)dim(\tau), DDIM performs better.
  • Runtime rises linearly with trajectory length, showing that DDIM generates samples more efficiently.
  • Quality requiring around 1,000 DDPM steps takes only 20–100 DDIM steps, ten to fifty times faster than DDPM.
 

5.2. Sample Consistency in DDIMs

  • DDIM generation is deterministic, and x0x_0 depends only on initial state xTx_T. Figure 5 compares trajectories for the same xTx_T; the same initialxTx_T preserves similar high-level features.
  • Thus, xTx_T alone is an informative latent image encoding.
  • Longer trajectories improve quality without significantly changing high-level features.
 

5.3. Interpolation in Deterministic Generative Processes

  • Since xTx_T encodes high-level features, the authors explore semantic interpolation, as in GANs, another implicit-model family. xTx_TInterpolation between two samples is possible.
  • This is not possible in the same way with DDPM's stochastic generation.
 

5.4. Reconstruction from Latent Space

  • DDIM's Euler-integrated ODE can encode x0x_0 into xTx_T and reconstruct in reverse. As in neural ODEs, larger trajectory length SS lowers reconstruction error. DDPM does not offer the same reconstruction process.
 

Review

DDIM substantially improves DDPM sampling speed. Moving from Markovian to non-Markovian and stochastic to deterministic processes also enables interpolation, encoding, and reconstruction.

 

References

https://arxiv.org/abs/2010.02502

https://happy-jihye.github.io/diffusion/diffusion-2/

 
 

Read next