LDM: High-Resolution Image Synthesis with Latent Diffusion Models

This post has turned into something close to a paper translation, but there are many interesting ideas: reducing computing demand, democratizing diffusion development, and limiting environmental harm. Perhaps the decision to release Stable Diffusion as open source connects with these aims.
 

LDM: High-Resolution Image Synthesis with Latent Diffusion Models

 
 

0. Abstract

LDM과 다른 Methods간 성능비교; 훨씬 선명하고 디테일하다
LDM compared with other methods: Much sharper and more detailed.

Diffusion models achieve state-of-the-art image and data generation. However, operating directly in pixel space requires substantial GPU resources and time. To retain quality and flexibility while training with limited resources, the authors use a powerful pretrained autoencoder's latent space. Training diffusion in this representation finds a near-optimal balance between complexity reduction and detail preservation.

Cross-attention makes conditioning on inputs such as text and bounding boxes powerful and flexible, while convolutional processing supports high resolution.

Latent diffusion models achieve state of the art in inpainting and class-conditioned synthesis, plus strong performance in text-to-image generation, unconditional generation, and super-resolution, with lower computational requirements than pixel-based diffusion.

 

1. Introduction

Image synthesis has advanced dramatically and demands considerable computation. Until recently, generating complex natural scenes at high resolution relied on scaling autoregressive transformers to billions of parameters. Adversarial training makes GANs difficult to scale to complex multimodal distributions, limiting variability.

Diffusion models built from denoising autoencoders deliver impressive synthesis and state-of-the-art class-conditioned generation and super-resolution. Their likelihood-based approach avoids GAN mode collapse and training instability while enabling parameter sharing. They can model highly complex natural-image distributions without enormous parameter counts.

 

Democratizing high-resolution image synthesis

Diffusion still requires substantial computing resources because training and evaluation operate in high-dimensional RGB space. Strong models can require 150–1,000 V100-days of training, with two consequences:

  1. Training needs resources accessible to only a small portion of the field and creates a large carbon footprint.
  1. Evaluating trained models is expensive in time and memory.

Improving accessibility requires reducing both training and sampling complexity. Maintaining performance while lowering demand is key to democratizing high-resolution synthesis.

 

Moving into latent space

The approach begins by analyzing pixel-space models. Figure 2 shows the rate–distortion trade-off: LDM improves efficiency by discarding imperceptible detail. The aim is a perceptually equivalent, computationally suitable space. Training has two stages:

  1. Train an autoencoder providing a lower-dimensional representation perceptually equivalent to the data. Train diffusion in that latent space, which scales better than the original spatial dimensions.
  1. Reduced complexity enables effective generation in latent space, followed by decoding through a single network pass. These are called latent diffusion models.

3. Method

3.1. Perceptual Image Compression (Pixel Space ↔ Latent Space)

The perceptual compressor is an autoencoder. For RGB image x∈RH×W×3x \in \mathbb{R}^{H \times W \times 3}, encoder E\mathcal{E} maps xx to latent representation z=E(x), z∈Rh×w×cz=\mathcal{E}(x), \ z \in \mathbb{R}^{h\times w \times c}, and decoder D\mathcal{D} reconstructs image x~=D(z)=D(E(x))\tilde{x} = \mathcal{D}(z) = \mathcal{D}(\mathcal{E}(x)). The encoder downsamples by factor f=H/h=W/wf=H/h = W/w; experiments examine factors f=2m, m∈Nf=2^m, \ m \in \mathcal{N}.

Two regularization approaches are tested to avoid high variance in the latent space.

  1. KL-reg: A small KL penalty towards a standard normal distribution over the learned latent, similar to VAE.
  1. VQ-reg: Uses a vector quantization layer within the decoder, like VQVAE but the quantization layer is absorbed by the decoder.

The compressor preserves a two-dimensional structure in latent space z=E(x)z=\mathcal{E}(x), retaining more detail in xx than earlier one-dimensional representations zz.

 

3.2. Latent Diffusion Models

Diffusion Model

Diffusion models learn data distribution p(x). They can be viewed as a weighted sequence of denoising autoencoders ϵθ(xt,t); t=1,…,T\epsilon_\theta(x_t, t); \ t= 1, …,T predicting the original image xx from noisy input xtx_t. The simplified objective is:

 

Generative Modeling of Latent Representations

A compressor comprising E\mathcal{E} and D\mathcal{D} gives access to a low-dimensional space better suited to likelihood-based generation: it focuses on important semantic information and trains efficiently in fewer dimensions.

The backbone ϵθ(○,t)\epsilon_\theta(\mathrm{○}, t) is a time-conditioned U-Net. Because the forward process is fixed, ztz_t is easily obtained through E\mathcal{E} during training. Samples from p(z)p(z) decode through D\mathcal{D} in one pass.

 

3.3. Conditioning Mechanisms

Like other generative models, diffusion models conditional distributions p(z∣y)p(z|y). Conditional denoiser ϵθ(zt,t,y)\epsilon_\theta(z_t, t, y) controls generation through input yy, such as text or semantic maps, and enables image-to-image translation. The U-Net uses cross-attention to accommodate different modalities.

A domain-specific encoder τθ\tau_\theta preprocesses conditions yy from modalities such as text and semantic maps, projecting yy into τθ(y)∈RM×dτ\tau_\theta(y) \in \mathbb{R^{\mathrm{M} \times d_\tau}} for the U-Net.

UNet 내부의 cross attention mechanism
Cross-attention inside the U-Net
  • QQ is the query and KK the key, representing the input embedding. Their dot product passes through softmax to produce weights, which combine with value VV, also an input embedding like KK. This is standard cross-attention.
  • φi(zt)∈RN×dτϵ\varphi_i(z_t) \in \mathbb{R^{\mathrm{N} \times d_\tau^\epsilon}} denotes the flattened U-Net representation of ϵθ\epsilon_\theta.
  • WV(i)∈Rd×dϵiW_V^{(i)} \in \mathbb{R^{\mathrm{d} \times \mathrm{d}_\epsilon^i}}, WQ(i)∈Rd×dτW_Q^{(i)} \in \mathbb{R^{\mathrm{d} \times \mathrm{d}_\tau}}, and WV(i)∈Rd×dϵiW_V^{(i)} \in \mathbb{R^{\mathrm{d} \times \mathrm{d}_\epsilon^i}} are trainable projection matrices.
 

The simplified conditional LDM objective is:

4. Experiments

4.1. On Perceptual Compression Tradeoffs

Experiments vary downsampling factor f∈{1,2,4,8,16,32}f \in \{1,2,4,8,16,32\}, naming models LDM-ff. LDM-1 is pixel-based diffusion. LDM-{4–16} balances efficiency and quality, with LDM-4 and LDM-8 optimal for high-quality results.

 

4.2. Image Generation with Latent Diffusion

  • Achieves state-of-the-art FID on CelebA-HQ.
  • Maintains low FID with relatively few parameters, around one billion.
 

4.3. Conditional Latent Diffusion

<Text-To-Image>

<Layout-To-Image>

<Class Conditional Image Geneartion>

 

4.4. Super-Resolution with Latent Diffusion

 

4.5. Inpainting with Latent Diffusion

 

5. Limitation & Societal Impact

Limitation

LDM reduces computation greatly, but sampling remains slower than GANs. Applications needing high pixel-level precision may be limited by autoencoder reconstruction. Although quality loss in LDM-4 is small, as Figure 1 shows, this can be a bottleneck. The authors suspect similar limitations already affect super-resolution.

 

Societal Impact

Media-generation models are a double-edged sword.

They enable creative applications, and reducing training and inference costs improves access and democratizes research. But it also makes manipulated data, misinformation, and spam easier to distribute. Deliberate image manipulation, or deepfakes, is a recurring concern that disproportionately affects women.

Generative models may expose training data, raising serious concerns when sensitive or personal data was collected without explicit consent. The extent of this in diffusion images is not fully understood. Deep learning also reproduces or amplifies existing biases. Diffusion covers distributions better than GAN approaches, but how much LDM misrepresents data remains an important research question.

For ethical discussion, see Ethical Considerations of Generative AI, Emily Denton, CVPR 2021.

 

6. Conclusion

The paper introduces LDM, a simple efficient approach improving diffusion training and sampling without sacrificing quality. Cross-attention supports broad conditional synthesis tasks without task-specific architectures, with performance comparable to state-of-the-art models.

 

Review

Overall, the paper prioritizes concepts over equations. Its state-of-the-art results across tasks and the visual quality are striking.

The figure below still confuses me a little: I do not fully understand how DDIM acts as a sampler here. I know it is used as a scheduler in Stable Diffusion, but wonder how its algorithm interacts with the architecture.

 
 

Reference

https://arxiv.org/abs/2112.10752

https://velog.io/@hewas1230/StableDiffusion

 

Read next