An Overview of Generative Models

A brief introduction to several approaches to image generation, followed by an introduction to diffusion models, which have recently attracted considerable attention.
 
 

1. Types of generative models

생성 모형의 종류
Types of generative models

(1) Generative Advarsarial Network(GAN)

Proposed by Ian Goodfellow in 2014, this framework trains a generator to create realistic images through minimax, adversarial training between a generator and a discriminator.

GAN Model의 구조
GAN architecture
  • Advantages: High fidelity
    • Can generate coherent images that resemble real photographs.
    • Offers flexibility in the choice of model architecture.
  • Disadvantages: Less diversity
    • Limited distribution coverage means that the model generates realistic images, but not a wide variety of them.
    • Training can be unstable.
 

(2) Variational Autoencoder(VAE)

VAE Model의 구조
VAE architecture
  • Advantage: Its probabilistic formulation allows more flexible computation of latent codes.
  • Disadvantage: Because it does not calculate density directly, its performance falls short of models such as PixelRNN and PixelCNN that model density explicitly.
 

(3) Likelihood Based Model

A flow model models a complex distribution starting from a simple one, such as a normal distribution, through invertible transformations implemented by a deep learning model. It is based on the change-of-variables theorem in probability theory and is trained with a log-likelihood objective.

flow model 구조
Flow model architecture

Autoregressive Model(DALL.E)

As in DALL·E (2021), image patches are tokenized through vector quantization into visual tokens, and a transformer decoder models relationships between those tokens. This adapts GPT, used for language generation in NLP, to another domain. The approach is used not only for images but also for language and speech generation. Its transformer-decoder architecture offers advantages for model scaling.

 

(4) Diffusion Model

A diffusion model is a deep generative model that creates data using two processes: a forward, or diffusion, process that gradually adds noise until the data becomes pure noise, and a reverse process that gradually reconstructs data from noise.

  • Advantages:
    • Uses a stationary training objective, unlike GANs.
    • Model scalability(CNN architecture)
    • High distribution coverage allows diverse images to be generated.
  • Disadvantages:
    • Generation is relatively slow because images are produced through a sequential reverse process.
    • Lower fidelity than GANs.
 

2. DPM

In deep learning, representing a complex real-world dataset as a probability distribution probability distribution is important. When seeking that distribution, tractability and flexibility are particularly important, but they trade off against each other. It is difficult to find a distribution that both fits complex data well and is easy to compute.

 
  • Tractability: A distribution, such as a Gaussian or Laplace distribution, that is easy to fit to data, analyze, and compute.
  • Flexibility: A distribution that can accommodate arbitrary complex data.
 
Diffusion Probability Model
  1. extreme flexibility in model structure
  1. exact sampling
  1. easy multiplication with other distributions, e.g. in order to compute a posterior, and
  1. the model log likelihood, and the probability of individual states, to be cheaply evaluate
 

Early Diffusion Probability Model(2015) work sought a distribution that was both flexible and tractable by learning a Markov chain that transforms a familiar distribution, such as a Gaussian, into the target data distribution through diffusion.

VAE trains both a network that encodes images and one that decodes images from latent codes. In contrast, Diffusion model keeps its image-encoding forward process fixed and trains only the image-decoding reverse process—a single network.

 

References


  1. https://lilianweng.github.io/posts/2021-07-11-diffusion-models/
  1. https://velog.io/@dongdori/Denoising-Diffusion-Probabilistic-Model2020
  1. https://happy-jihye.github.io/diffusion/diffusion-1/
 
 

Read next