LDM: High-Resolution Image Synthesis with Latent Diffusion Models
This post has turned into something close to a paper translation, but there are many interesting ideas: reducing computing demand, democratizing diffusion development, and limiting environmental harm. Perhaps the decision to release Stable Diffusion as open source connects with these aims.
LDM: High-Resolution Image Synthesis with Latent Diffusion Models
0. Abstract

Diffusion models achieve state-of-the-art image and data generation. However, operating directly in pixel space requires substantial GPU resources and time. To retain quality and flexibility while training with limited resources, the authors use a powerful pretrained autoencoder's latent space. Training diffusion in this representation finds a near-optimal balance between complexity reduction and detail preservation.
Cross-attention makes conditioning on inputs such as text and bounding boxes powerful and flexible, while convolutional processing supports high resolution.
Latent diffusion models achieve state of the art in inpainting and class-conditioned synthesis, plus strong performance in text-to-image generation, unconditional generation, and super-resolution, with lower computational requirements than pixel-based diffusion.
1. Introduction
Image synthesis has advanced dramatically and demands considerable computation. Until recently, generating complex natural scenes at high resolution relied on scaling autoregressive transformers to billions of parameters. Adversarial training makes GANs difficult to scale to complex multimodal distributions, limiting variability.
Diffusion models built from denoising autoencoders deliver impressive synthesis and state-of-the-art class-conditioned generation and super-resolution. Their likelihood-based approach avoids GAN mode collapse and training instability while enabling parameter sharing. They can model highly complex natural-image distributions without enormous parameter counts.
Democratizing high-resolution image synthesis
Diffusion still requires substantial computing resources because training and evaluation operate in high-dimensional RGB space. Strong models can require 150–1,000 V100-days of training, with two consequences:
- Training needs resources accessible to only a small portion of the field and creates a large carbon footprint.
- Evaluating trained models is expensive in time and memory.
Improving accessibility requires reducing both training and sampling complexity. Maintaining performance while lowering demand is key to democratizing high-resolution synthesis.
Moving into latent space

The approach begins by analyzing pixel-space models. Figure 2 shows the rate–distortion trade-off: LDM improves efficiency by discarding imperceptible detail. The aim is a perceptually equivalent, computationally suitable space. Training has two stages:
- Train an autoencoder providing a lower-dimensional representation perceptually equivalent to the data. Train diffusion in that latent space, which scales better than the original spatial dimensions.
- Reduced complexity enables effective generation in latent space, followed by decoding through a single network pass. These are called latent diffusion models.
3. Method

3.1. Perceptual Image Compression (Pixel Space ↔ Latent Space)
The perceptual compressor is an autoencoder. For RGB image , encoder maps to latent representation , and decoder reconstructs image . The encoder downsamples by factor ; experiments examine factors .
Two regularization approaches are tested to avoid high variance in the latent space.
- KL-reg: A small KL penalty towards a standard normal distribution over the learned latent, similar to VAE.
- VQ-reg: Uses a vector quantization layer within the decoder, like VQVAE but the quantization layer is absorbed by the decoder.
The compressor preserves a two-dimensional structure in latent space , retaining more detail in than earlier one-dimensional representations .
3.2. Latent Diffusion Models
Diffusion Model
Diffusion models learn data distribution p(x). They can be viewed as a weighted sequence of denoising autoencoders predicting the original image from noisy input . The simplified objective is:

Generative Modeling of Latent Representations
A compressor comprising and gives access to a low-dimensional space better suited to likelihood-based generation: it focuses on important semantic information and trains efficiently in fewer dimensions.

The backbone is a time-conditioned U-Net. Because the forward process is fixed, is easily obtained through during training. Samples from decode through in one pass.
3.3. Conditioning Mechanisms
Like other generative models, diffusion models conditional distributions . Conditional denoiser controls generation through input , such as text or semantic maps, and enables image-to-image translation. The U-Net uses cross-attention to accommodate different modalities.
A domain-specific encoder preprocesses conditions from modalities such as text and semantic maps, projecting into for the U-Net.

- is the query and the key, representing the input embedding. Their dot product passes through softmax to produce weights, which combine with value , also an input embedding like . This is standard cross-attention.
- denotes the flattened U-Net representation of .
- , , and are trainable projection matrices.
The simplified conditional LDM objective is:

4. Experiments
4.1. On Perceptual Compression Tradeoffs
Experiments vary downsampling factor , naming models LDM-. LDM-1 is pixel-based diffusion. LDM-{4–16} balances efficiency and quality, with LDM-4 and LDM-8 optimal for high-quality results.


4.2. Image Generation with Latent Diffusion



- Achieves state-of-the-art FID on CelebA-HQ.
- Maintains low FID with relatively few parameters, around one billion.
4.3. Conditional Latent Diffusion
<Text-To-Image>

<Layout-To-Image>

<Class Conditional Image Geneartion>

4.4. Super-Resolution with Latent Diffusion


4.5. Inpainting with Latent Diffusion


5. Limitation & Societal Impact
Limitation
LDM reduces computation greatly, but sampling remains slower than GANs. Applications needing high pixel-level precision may be limited by autoencoder reconstruction. Although quality loss in LDM-4 is small, as Figure 1 shows, this can be a bottleneck. The authors suspect similar limitations already affect super-resolution.
Societal Impact
Media-generation models are a double-edged sword.
They enable creative applications, and reducing training and inference costs improves access and democratizes research. But it also makes manipulated data, misinformation, and spam easier to distribute. Deliberate image manipulation, or deepfakes, is a recurring concern that disproportionately affects women.
Generative models may expose training data, raising serious concerns when sensitive or personal data was collected without explicit consent. The extent of this in diffusion images is not fully understood. Deep learning also reproduces or amplifies existing biases. Diffusion covers distributions better than GAN approaches, but how much LDM misrepresents data remains an important research question.
For ethical discussion, see Ethical Considerations of Generative AI, Emily Denton, CVPR 2021.
6. Conclusion
The paper introduces LDM, a simple efficient approach improving diffusion training and sampling without sacrificing quality. Cross-attention supports broad conditional synthesis tasks without task-specific architectures, with performance comparable to state-of-the-art models.
Review
Overall, the paper prioritizes concepts over equations. Its state-of-the-art results across tasks and the visual quality are striking.
The figure below still confuses me a little: I do not fully understand how DDIM acts as a sampler here. I know it is used as a scheduler in Stable Diffusion, but wonder how its algorithm interacts with the architecture.

Reference
https://arxiv.org/abs/2112.10752
https://velog.io/@hewas1230/StableDiffusion