StarGAN v2 (2020)

Github: https://github.com/clovaai/stargan-v2

Paper: https://arxiv.org/abs/1912.01865

 

Contents

StarGAN v2: Diverse Image Synthesis for Multiple Domains

Yunjey Choi, Youngjung Uh, Jaejun Yoo, Jung-Woo Ha

A good image-to-image translation model should learn a mapping between different visual domains while satisfying the following properties: 1) diversity of generated images and 2) scalability over multiple domains. Existing methods address either of the issues, having limited diversity or multiple models for all domains. We propose StarGAN v2, a single framework that tackles both and shows significantly improved results over the baselines. Experiments on CelebA-HQ and a new animal faces dataset (AFHQ) validate our superiority in terms of visual quality, diversity, and scalability. To better assess image-to-image translation models, we release AFHQ, high-quality animal faces with large inter- and intra-domain differences. The code, pretrained models, and dataset can be found at this https URL.
 

Abstract

  • A good image-to-image translation model should offer:
    • 1. Diversity in generated images.

      2. Scalability across many domains.

  • StarGAN v2 satisfies both in one framework. Experiments on CelebA-HQ and AFHQ validate its superior visual quality, diversity, and scalability.

1. Introduction

Terminology

  • Domain: set of images that can be grouped as a visually distinctive category.
  • Style: each image has a unique appearance
    • For example, domain can represent gender, while style includes makeup, beards, and hairstyles, as in the CelebA-HQ examples above.
  • An ideal translation method should synthesize different styles within each domain.
    • But arbitrary styles and domains are so numerous that this is difficult.

 
 

A brief look at an accessible explanation of GANs

https://www.notion.so/StarGAN-v2-53805dc3615a49aea9de987dd93b59b6#6a79a22ba77344f798819bbb980e75f3
 

Earlier approaches

  • Many methods inject low-dimensional latent codes into generators for different styles—for example, a randomly sampled 100-dimensional uniform or Gaussian vector.
  • Domain-specific decoders treat those vectors as recipes for generation.
  • However, these methods consider only mappings between two domains.
    • They do not scale with the number of domains.

      Learning K domains requires K(K−1) generators.

      This limits practical usefulness.

       

What is different about StarGAN?

  • Unified frameworks address scalability; StarGAN was among the first.
    • One generator learns all possible domain mappings.

  • It takes a domain label as an additional input and learns the corresponding transformation.
  • But it learns only deterministic mappings, not the distribution's multimodal nature.
    • Multimodal: Our experiences are multimodal—seeing, hearing, touch, smell, and taste. A modality is a way something happens or is experienced; using these together requires multimodal characterization.

      Multimodal data: Information in different forms with distinguishable characteristics.

  • The limitation comes from determining each domain through a fixed label.
  • A fixed input such as a one-hot label inevitably leads to the same transformation for each domain.
 

StarGAN v2: The solution

  • Starting from StarGAN's scalability solution, replace domain labels with newly designed domain-specific style codes to capture multimodality.
    • Style codes represent distinct styles within a particular domain!

  • Two modules enable this:
      1. Mapping network: Transforms random Gaussian noise into a style code.
      1. Style encoder: Extracts a style code from a reference image.
  • These codes teach the generator to synthesize varied images across multiple domains.
 
  • Component studies show that style codes are effective; see Section 3.1.
  • The method scales across domains and improves visual quality and diversity over earlier methods.
 

2. StarGAN v2

2.1 Proposed framework

  • Let X and Y be the image set and possible domains.
    • → X = sets of images, Y = sets of possible domain

  • For image x and target domain y, a single generator G should produce diverse corresponding images.
    • Given an image x and a blond-hair domain, G should generate varied blond versions of x.

      Perhaps replacing labels with style codes also captures complex combinations—female, blond, Western appearance, and so on—in my interpretation of multimodal data.

  • Generate domain-specific vectors within learned style spaces, and train G to reflect them.
  • Figure 2 gives the framework overview, comprising four modules.
 
(a) Generator
  • Generator G transforms x into G(x, s), reflecting domain-specific code s from either mapping network F or style encoder E.
  • Adaptive instance normalization (AdaIN) injects s into G.
  • Since s represents the style of domain y, G does not need y separately. It can synthesize images across all domains!
    • Using style codes instead of domain labels preserves scalability.

       
(b) Mapping network
  • Given latent z and domain y, mapping network F produces style code s=Fy(z)s = F_y(z).
  • F is an MLP with multiple output branches covering all domains.
  • F produces diverse codes by randomly sampling z∈Zz \in Z and y∈Yy \in Y.
  • This multitask architecture learns style representations efficiently and effectively across domains.
    • —>

 
(c) Style encoder
  • For image x and its domain y, encoder E extracts style code s=Ey(s)s = E_y(s).
  • Like F, E benefits from multitask learning and produces diverse codes from different references.
  • G can therefore synthesize outputs reflecting reference x's style s!
 
(d) Discriminator
  • D is a multitask discriminator with several output branches.
  • Each branch Dy D_y classifies x as a real image in domain y or fake G(x,s)G(x,s).
 

2.2 Training Objectives

Training uses the following objectives for image x∈Xx \in X and original domain y∈Yy \in Y.

 
Adversarial objective
  • Randomly sample latent z∈Zz \in Z and target domain y~∈Y\tilde{y} \in Y, generating target code s~=Fy~(z)\tilde{s} = F_{\tilde{y}}(z).
  • G learns to produce G(x,s~)G(x, \tilde{s}) from x and s~\tilde{s} through adversarial loss.
    •  
       
  • Dy(.)D_{y}(.) is the discriminator output for domain y.
  • F learns code s~ \tilde{s} for target domain y~\tilde{y}; GG uses s~\tilde{s} to generate G(x,s~)G(x,\tilde{s}) indistinguishable from real images in y~\tilde{y}.
 
Style reconstruction
  • Style-reconstruction loss encourages G to use code s~\tilde{s}.
    •  
       
  • This resembles earlier image-to-latent approaches using multiple encoders.
  • The difference is training only a single encoder E for multiple domains!
    • This connects to the advantages of style codes.

  • At test time, E lets G transform an input to reflect a reference's style.
 
Style diversification
  • Diversity-sensitive loss explicitly regularizes G to produce more diverse images.
 
  • F's target codes s~1\tilde{s}_1 and s~2\tilde{s}_2 come from S~i=Fy~(zi)\tilde{S}_i = F_{\tilde{y}}(z_i) for i∈{1,2}i \in \{1,2 \}.
  • Maximizing this term encourages exploration of image space and meaningful style features.
  • In earlier formulations, a tiny denominator ∥z1−z2∥1\parallel z_1 - z_2 \parallel_1 could make loss explode and destabilize training.
  • The paper removes that denominator, keeping the intuition while improving stability.
 
Preserving source characteristics
  • Cycle-consistency loss preserves domain-invariant characteristics, such as pose, in generated image G(x,s~)G(x,\tilde{s}).
 
  • s^=Ey(x)\hat{s} = E_y(x) is x's estimated style code, and y is its original domain.
  • Including estimated code s^\hat s when reconstructing x lets G change style faithfully while preserving original characteristics.
    • So the source face stays recognizable because ofs^\hat s!

 
Full objective
  • λ\lambda terms are the loss hyperparameters.
  • Important: The same objectives also apply when references and the style encoder generate codes instead of latents and the mapping network. More details are in Appendix B.

This seems to concern LstyL_{sty}, originally the objective for mapping network F converting latents into styles.

Does the same objective apply to encoder E converting references into styles?

 

3. Experiments

  • Experiments examine components and compare three leading baselines.
  • All use unseen data.
 
Baselines
  • MUNIT, DRIT, and MSGAN perform multimodal mapping between two domains.
  • For multi-domain comparison, train them separately on each image-domain pair.
  • Also compare StarGAN v1, which uses one generator for multiple domains.
 
Datasets
  • Evaluate on CelebA-HQ and AFHQ.
  • Divide CelebA-HQ into male and female domains.
  • Use no additional information beyond domain labels, allowing unsupervised style learning.
  • Resize images to 256×256 for fair comparison, the baselines' maximum resolution.
 
Evaluation metrics
  • There are two evaluation metrics.
    • 1. visual quality —> FID(Frechet Inception distance): distance between two distributions of real and generated images (lower is better)

      Question: Does FID include visual quality?

      1. diversity of generated images —> LPIPS(Learned Perceptual Image Patch Similarity): (higher lis better)
  • The table adds v2 components to StarGAN v1 one at a time.
  • (A): StarGAN v1 mainly transfers makeup, producing local changes.
  • (B): Replacing the ACGAN discriminator with a multitask discriminator enables global structure changes.
  • (C): R1 regularization and AdaIN improve training stability.
    • A–C cannot produce multiple outputs for one input and target domain, so LPIPS diversity cannot be measured from a single output.

       
  • (D): Directly injecting latent z into G might add diversity, but in multi-domain experiments it fails to learn meaningful styles or expected diversity.
 
  • Therefore, the authors hypothesize:
      1. Latent codes cannot distinguish domains.
      1. Latent-reconstruction loss models shared styles rather than domain-specific ones.
       
  • Instead, F maps z to domain-specific s, which is injected into G.
  • Style-reconstruction loss is also introduced.
  • Each mapping-network branch corresponds to a particular domain, removing ambiguity in style codes.
  • Unlike latent-reconstruction loss for F, style-reconstruction loss for E encourages domain-specific styles—apparently using reference-derived codes for generation.
 

3.2 Comparison on diverse image synthesis

  • Compare two settings:
      1. Latent-guided synthesis: Results remain intact where baselines show artifacts and generally look better.
      1. Reference-guided synthesis: Extracted codes capture distinctive styles, unlike earlier methods that mainly transfer color distributions.
  • Human-evaluation
    • Human preference voting also favors this method.
 

4. Discussion

First

Multi-head mapping networks and style encoders generate codes separately for each domain.

The generator can focus solely on using those codes.

 
Second

Baselines assume a fixed Gaussian style distribution.

Following StyleGAN, v2 generates in a learned transformed style space.

 
  • Reference-guided synthesis also works well on FFHQ.

5. Related work

  • StyleGAN uses a nonlinear mapping from input latents into style space.
    • But it does not accept images, so transforming a real image is not straightforward.

  • Both latent- and reference-guided synthesis train on coarsely labeled datasets.
 
결국 reference guided synthesis로 갈텐데, 이때 우리는 스타일 합성을 하는 것이지, 헤어 합성이 아니다. —> 초기에는 스타일 합성 기능을 제공하는 것으로,,? —> 대신 이를 위해 동양인 dataset에 대한 fine tuning이 필요함. -> celebA-HQ 30000장, 1024x1024 1. 초기 launching시에 우리가 제공하는 스타일들을 보여주고, 10개정도 scrap할 수 있게(왓챠나 넷플릭스 처럼)함. 2. 우리는 AI 스타일 합성을 제공하며, 이는 해당 사진의 헤어, 메이크업 등을 포함함 3. 따라서 우리는 스타일 전반에 대한 설명 또한 제공(메이크업 정보, 헤어 정보 등)
 

6. Conclusion

  • StarGAN v2 maps one domain into diverse images of a target domain and supports multiple targets.
 

Appendix

B. Training details

  • batch size: 8
  • iteration: 100,000(100K)
  • One Tesla V100 GPU, three days.
  • λsty=1,  λds=1,  λcyc=1\lambda_{sty} = 1, \; \lambda_{ds} = 1, \; \lambda_{cyc} = 1 for CelebA-HQ
  • non-saturating adversarial loss + R1 regularization γ=1\gamma = 1
  • Adam optimizer with β1=0, β2=0.99\beta_1 = 0, \: \beta_2 = 0.99
  • Learning rates for G,D,and  E=10−4G, D, and \; E = 10^{-4}
  • Learning rates for F=10−6F = 10^{-6}
  • initialize the weights = He initialization
  • set biases to zero —> except for biases associated with scaling vector of AdaIN(set to one)

E. Network architecture

Generator
 
 
Mapping Network

K output branches correspond to the number of domains.

The first four fully connected layers are shared; the final four are domain-specific.

Dimensions of each component
  • latent code: 16
  • hidden layer: 512
  • style code: 64

Sample latents from a standard Gaussian.

No pixel normalization is applied to latents, since it did not improve performance.

Feature normalization was also ineffective.

 
Style encoder

—> CNN with K output branch

Six pre-activation residual blocks are shared across domains.

Dimension information
  • Output: 64 x K
 
Discriminator
  • The multitask discriminator uses the architecture shown in the table above.
  • K fully connected outputs classify real versus fake for each domain.
    • Dimension information

    • D = 1 (real/fake)
  • The multitask discriminator outperforms other conditional discriminators.
 

Fine-tuning notes

 
  • Eight Azure V100 GPUs

Original estimate: 30,000 training images take three days, or 72 hours, on one GPU; eight GPUs imply nine hours. 2,725 × 9 ≈ 24,000 won.

  • V100 else
    •  

Read next