DDIM: Denoising Diffusion Implicit Models
DDPM teaches a generative Markov chain that removes noise gradually. Sequential sampling needs many steps and is much slower than single-pass GANs. DDIM generalizes that process into a non-Markovian one.
DDIM: Denoising Diffusion ImplicitModels
0. Abstract
DDPM generates high-quality images without adversarial training, but requires a long Markov chain and many sampling steps.
Denoising diffusion implicit models (DDIM) accelerate sampling by generalizing DDPM to non-Markovian diffusion. This enables a more deterministic generative process and high-quality samples ten to fifty times faster.
1. Introduction
Deep generative models have demonstrated high-quality generation across image domains. GANs outperform likelihood-based VAEs and flows in quality, but stable training restricts architectures and optimization, and they fail to cover all distribution modes.
DDPM and noise-conditional score networks produce GAN-quality samples without adversarial training by denoising through a generative Markov chain. Many iterations make them much slower than single-pass GANs.
To narrow this efficiency gap, DDIM uses the same training objective as DDPM but generalizes the forward process to non-Markovian chains. This can improve sampling efficiency ten to one hundred times with only a small quality loss.
DDIM also offers stronger consistency: samples starting from the same latent have similar high-level features. This enables semantically meaningful interpolation.
2. Background
If all conditionals are Gaussian with learnable means and fixed variances, DDPM's objective simplifies as follows.

Let be a collection of functions. Each timestep-indexed depends on trainable parameters and positive-coefficient vector , whose coefficients depend on . DDPM optimizes objective for generation performance. To generate , sample from prior , then repeatedly sample .
Forward-process length T is an important hyperparameter. Larger T makes reverse transitions closer to Gaussian and improves their approximation. DDPM therefore uses large T, such as 1,000. But these T steps are sequential, not parallel, making sampling slower than other generative models.
3. Variational Inference for Non-Markovian Forward Process

Reconsidering inference, the authors find that objective depends only on marginal distributions , not directly on joint distribution . Many joints share those marginals, allowing a new non-Markovian process, shown on Figure 1's right. They establish this for Gaussian cases.
Definitions of joint and marginal distributions
3.1. Non-Markovian Forward Process

To satisfy these equations, the forward inference mean must be:
The mean ensures at every t, so the joint distribution has the desired marginals. Bayes' rule gives the following forward process.
Unlike DDPM, DDIM's forward process is no longer Markovian, because depends on both and . controls forward-process stochasticity. As it approaches zero, becomes fixed and deterministic.
3.2. Generative Process and Unified Variational Inference Objective
Next define learnable . uses known reverse conditional to sample from noisy using . Intuitively, for observation , first predict corresponding , then sample through .
Given and , Equation 4 gives . Model predicts from without knowing . Rearranging for yields 's prediction for , the denoised observation. Fixed prior then defines the generative process.
(1) denoised observation
- Expressing Equation 4 in terms of
- Predicting with
(2) generative process
- In , denoised observation replaces .
- Gaussian noise is added when to ensure a generative process.
(3) variational inference objective

- Different choices of yield different models.
- Setting gives the DDIM objective.
- Setting gives the DDPM objective.
(4) Theorem 1.
- If , the objective equals DDPM's variational lower bound.
- Theorem 1 shows that is a form of ; has its optimum when ().

4. Sampling from Generalized Generative Process
Since gives the optimum, use as the objective.
Choosing yields Markovian or non-Markovian processes. Regardless of , the parameters to learn are .
- Thus, pretrained DDPM parameters can also be used in DDIM's generative process.
- Rather than introducing a new training method, DDIM generalizes diffusion to a non-Markovian chain and introduces a faster sampling method.
- A common approach is training with DDPM and sampling with DDIM, combining a strong model with fast generation.
4.1. Denoising Diffusion Implicit Model
In , the equation below generates sample from . is standard Gaussian noise independent of , and defines the setting. Different yield different generative processes, but shared means no retraining.

- forward process = “Markovian” = DDPM
- The forward process becomes deterministic with respect to and , except .
- becomes zero, making the model an implicit probabilistic model.
- The sample-generation process from to becomes fixed, allowing faster sampling.
- This is DDIM: An implicit probabilistic model trained with the DDPM objective.
- The forward process is no longer diffusion in this case.
4.2. Accelerated Generation Processes
Previously, generation approximated the reverse of a -step forward process, requiring sampling steps. But denoising objective , with fixed, no longer depends on that forward trajectory. We can consider fewer than steps, accelerating generation without training another model.
Suppose the forward process is defined not over all latents , but a subset. is an increasing subsequence of of length S. Define its transitions so matches the marginals.

Generation samples in reverse. This is the sampling trajectory. If the trajectory is shorter than , computational efficiency improves substantially.
Small changes to Equation 12 yield a faster process applicable to DDPM and DDIM. Subsampled trajectories generate images faster, but if a model trains only at selected forward steps, generation can use only those steps; broader timestep training is needed.
- Training at many steps as in DDPM is more effective than training only selected steps in this interpretation of DDIM.
- Hence DDPM training with DDIM sampling is common.
4.3. Relevance to Neural ODEs
What is an ordinary differential equation?
Rewriting DDIM iterations as Equation 12 makes their similarity to Euler integration for ordinary differential equations clearer.

To derive the ODE, reparameterize () with and () with . When , Equation 13 becomes Euler's method for the ODE below.

With sufficiently fine discretization, reversing Equation 14's ODE enables encoding, reversing generation: → .
5. Experiments
DDIM needs far fewer iterations than DDPM, giving ten- to one-hundred-fold speedups. With initial latent fixed, different trajectories preserve high-level features, and latent interpolation becomes possible.
DDIM can encode samples and reconstruct images from latents, unlike DDPM's stochastic process. Each dataset uses the same trained model with and objective . Only how samples are produced changes.
Experiments vary subsequences of and variance hyperparameter . is directly controllable.

- If : DDPM.
- If : DDIM.
5.1. Sample Quality and Efficiency

- Figure 3 reports FID for models trained on CIFAR-10 and CelebA.
- Quality improves as increases, trading off against computation. For the same , DDIM performs better.

- Runtime rises linearly with trajectory length, showing that DDIM generates samples more efficiently.
- Quality requiring around 1,000 DDPM steps takes only 20–100 DDIM steps, ten to fifty times faster than DDPM.
5.2. Sample Consistency in DDIMs

- DDIM generation is deterministic, and depends only on initial state . Figure 5 compares trajectories for the same ; the same initial preserves similar high-level features.
- Thus, alone is an informative latent image encoding.
- Longer trajectories improve quality without significantly changing high-level features.
5.3. Interpolation in Deterministic Generative Processes

- Since encodes high-level features, the authors explore semantic interpolation, as in GANs, another implicit-model family. Interpolation between two samples is possible.
- This is not possible in the same way with DDPM's stochastic generation.
5.4. Reconstruction from Latent Space

- DDIM's Euler-integrated ODE can encode into and reconstruct in reverse. As in neural ODEs, larger trajectory length lowers reconstruction error. DDPM does not offer the same reconstruction process.
Review
DDIM substantially improves DDPM sampling speed. Moving from Markovian to non-Markovian and stochastic to deterministic processes also enables interpolation, encoding, and reconstruction.
References
https://arxiv.org/abs/2010.02502
https://happy-jihye.github.io/diffusion/diffusion-2/



