DreamBooth: Fine-Tuning Text-to-Image Diffusion Models for Subject-Driven Generation

DreamBooth: Fine Tuning Text-to-Image Diffusion Models for Subject-Driven Generation

The GitHub link is a Stable Diffusion adaptation.

 
 

0. Abstact

Large text-to-image models generate high-quality images but cannot reliably reproduce subjects from a reference set. This paper fine-tunes them on a few subject images, binding a unique identifier to the subject. A class-specific prior-preservation loss leverages embedded semantic knowledge, enabling new scenes, poses, compositions, and lighting absent from the references.

1. Introduction

Synthesizing imaginary scenes has been challenging, but large text-to-image models have made remarkable progress. Their semantic priors link words such as 'dog' with many instances. Yet they cannot preserve a particular subject.

Their output domain is fundamentally limited. Even detailed descriptions produce different appearances, since text–object pairings are inherently diverse, as Figure 2 illustrates.

The paper proposes personalization by expanding the model's language–vision dictionary. Given three to five subject photos, its objective embeds that subject in the output domain, letting users generate it with a bound unique identifier.

 

Two new techniques are introduced:

  1. Represent a subject with a rare-token identifier.
  1. Fine-tune a diffusion-based text-to-image framework.
 

The framework has two stages—low-resolution generation followed by super-resolution—and fine-tuning likewise has two stages.

  1. Fine-tune low-resolution text-to-image generation with prompts containing a unique identifier, such as 'A [V] dog.'
      • Apply class-specific prior-preservation loss to prevent overfitting and language drift.
  1. Fine-tune super-resolution.
      • Preserve even small subject details with high fidelity.
 

The contributions are:

  1. Defining subject-driven generation: Given a few photos, generate the subject in varied contexts while preserving its key features with high fidelity.
  1. Proposing a few-shot fine-tuning technique for text-to-image diffusion that retains existing semantic knowledge.
 

2. Related Work

Related work considers whether subject-preserving novel generation could be solved without fine-tuning. Here is a brief summary of those alternatives.

 

(1) Compositing objects into scenes

Cloning a subject onto a new background cannot naturally adapt its lighting, interactions with objects, or shadows.

 

(2) Text-driven editing

CLIP-based GAN editing works on curated domains such as faces but struggles with diverse subject types. It handles global properties and local edits, not generating a specific subject in a new context.

 

(3) Text-to-image synthesis

Imagen, DALL·E 2, Parti, and CogView2 perform strongly but text alone offers limited fine control. Masks restrict areas, yet do not synthesize a subject in a new context.

 

(4) Diffusion-model inversion

This controls input latents, similar to manipulating W, W+, or S spaces in my GAN face-editing research. For diffusion, one must find a noise map or condition vector producing the desired image. It is a possible future solution, but currently limited by textual expressiveness and the pretrained model's output domain. See the papers below.

 

(5) Personalization: Fine-tuning for a specific domain

Personalization is common in recommendation and language models. GAN-based MyStyle fine-tunes for one facial identity, requiring around 100 photos. DreamBooth needs only three to five across varied subjects, including animals.

DreamBooth belongs to personalization!

 

3. Preliminaries

(1) Cascaded Text-to-Image Diffusion Models

Conditional diffusion model x^θ\hat{x}_\theta trains with squared error to denoise noisy image zt:=αtx+σtϵz_t:=α_tx+\sigma_t\epsilon, as follows.

xx is ground truth, cc a condition vector from text or another source, and ϵ∼N(0,I)\epsilon \sim \mathcal{N}(0, \mathrm{I}) noise. αt\alpha_t, σt\sigma_t, and wtw_t control the noise schedule and sample quality; time t follows t∼U([0,1])t \sim U([0,1]). Inference repeatedly denoises zt∼N(0,I)z_t \sim \mathcal{N}(0,\mathrm{I}) using deterministic DDIM or stochastic ancestral sampling. x0t^:=x^θ(zt,c)\hat{x_0^t}:=\hat{x}_\theta(z_t,c) predicts x.

State-of-the-art systems use cascaded diffusion: a 64×64 base text-to-image model followed by text-conditioned super-resolution from 64×\times64 to 256×\times256, then 256×\times256 to 1024×\times1024.

 

(2) Vocabulary Encoding

Text conditioning is crucial for visual and semantic fidelity. Earlier models use CLIP text embeddings mapped into image embeddings by a learned prior, or pretrained T5-XXL.

This paper uses T5-XXL to embed tokenized prompts. Vocabulary encoding is important preprocessing. Prompt PP first passes through tokenizer ff before becoming condition embedding cc. The tokenizer is SentencePiece.

Tokenizer ff maps PP into fixed-length vector f(P)f(P). Language model Γ\Gamma converts it to c:=Γ(f(P))c:=\Gamma(f(P)), and diffusion is conditioned on cc.

4. Method

The goal is new subject images with altered backgrounds, colors, species, shapes, poses, expressions, media, or other semantics. The high-level method follows.

First, embed the subject in the output domain and bind it to a unique identifier. Fine-tuning on a small set risks overfitting and language drift: forgetting other subjects in the same class and losing knowledge of its natural variations.

An autogenous class-specific prior-preservation loss reduces overfitting and drift by preserving generation of varied instances of the same class.

Super-resolution also needs fine-tuning to preserve detail. Naively training it to reproduce the instance misses important details. The paper provides training and testing insights for stronger preservation and recontextualization. Figure 4 shows the process; pretrained Imagen is the base.

 

4.1. Representing the Subject with a Rare-token Identifier

(1) Designing Prompts for Few-Shot Personalization

The goal is a new key–value pair in the dictionary, through few-shot fine-tuning. How should this be supervised? Human prompts and online data normally require detailed image descriptions, which are costly, variable, and subjective.

The paper instead labels images 'a [identifier] [class noun]'. The identifier is unique to the subject; the noun is a broad class, such as cat, dog, or watch, obtained from a classifier. Omitting it slows training and hurts performance. This labeling enables new poses and contexts.

Using existing words such as 'unique' or 'special' as identifiers requires disentangling their old meaning and attaching a new one, slowing learning and reducing performance. Using 'blue' for a gray subject, for example, mixes blue and gray in generated images.

An identifier with weak priors in both the language and diffusion models is needed. Random strings such as 'xxy5syt00' may appear literally in images or evoke related concepts, performing poorly like existing words.

 

(2) Rare-token Identifiers

The approach finds relatively rare vocabulary tokens and inverts them into text. Rare-token lookup gives identifier f(V^)f(\hat{V}); detokenizing f(V^)f(\hat{V}) yields text V^\hat{V}. Three or fewer Unicode characters without spaces work best. For example, V^\hat{V}, sks is vocabulary index 9,061 of 9,258.

 

4.2. Class-specific Prior Preserve Loss

(1) Few-shot Personalization of a Diffusion Model

Fine-tuning image set Xs:={xsi;i∈{0,…,N}}\mathcal{X}_s := \{x_s^i;i \in \{ 0, …, N \} \} with common condition vector csc_s from 'a [identifier] [class noun]' and original loss Ex,c,ϵ,t[wt∣∣xθ^(αtx+σtϵ,c)−x∣∣22]\mathbb{E}_{x,c,\epsilon,t}[w_t||\hat{x\theta}(\alpha_tx+\sigma_t\epsilon,c) -x||_2^2] causes overfitting and language drift.

 

(2) Issue-1: Overffiting

With few images, fine-tuning can overfit their appearance and context, as in Figure 12's top row. Regularization or selective layer tuning could help, but which layers should be frozen to balance fidelity and semantic flexibility is unclear. Empirically, tuning all layers works best but causes drift.

 

(3) Issue-2: Language Drift

Language-model research describes drift as gradually forgetting syntax and semantics during task-specific tuning. Here, after tuning on a few subject images, diffusion forgets the class prior and cannot generate other instances of that class, as in Figure 13's middle row.

 

(4) Prior-Preservation Loss

Prior-preservation loss supervises with samples generated by the original model, retaining its prior during few-shot tuning. A frozen model's ancestral sampler generates xpr=x^(zt1,cpr)x_{pr} = \hat{x}(z_{t_1}, c_{pr}) using initial noise zt1∼N(0,I)z_{t_1} \sim \mathcal{N}(0, \mathrm{I}) and condition cpr:=Γ(f(”a [class noun]”))c_{pr} := \Gamma(f(”a \ [class \ noun]”)).

Ex,c,ϵ,ϵ′,t[wt∣∣xθ^(αtx+σtϵ,c)−x∣∣22 +λwt′∣∣x^θ(αt′xpr+σt′ϵ′,cpr)∣∣22]\mathbb{E}_{x,c,\epsilon,\epsilon', t}[w_t||\hat{x\theta}(\alpha_tx+\sigma_t\epsilon,c) -x||_2^2 \ + \lambda w_{t'}|| \hat{x}_\theta(\alpha_{t'}x_{pr} + \sigma_{t'}\epsilon', c_pr{}) ||_2^2 ]

λ\lambda weights the preservation term. Figure 4 shows training. Despite its simplicity, the loss effectively addresses both problems. Learning rate 10−510^{-5} and roughly ∼200\sim200 epochs work best, taking around fifteen minutes on one TPU v4.

 

4.3. Personalized Instance-Specific Super-Resolution

The base model controls major visual semantics; super-resolution preserves details and photorealism. Without SR tuning, new subject images can contain artifacts, incorrect features, or lost detail, as in Figure 14's bottom row. Tuning 64×64 → 256×256 is essential; tuning 256×256 → 1024×1024 helps fine detail.

5. Experiments

5.1. Applications

(1) Recontextualization

Given model xθ^\hat{x_\theta}, prompts containing the identifier and class noun generate new subject images. Recontextualization usually uses 'a [V] [class noun] [context description]'.

(2) Art Renditions

Prompts such as 'a painting of a [V] [class noun] in the style of [famous painter]' or the equivalent for a sculpture produce artistic interpretations. This differs from transferring another image's style onto a fixed scene: it preserves identity and details while changing scene semantics, including poses and settings absent from the references.

(3) Expression Manipulation

Examples of generating new expressions

(4) Novel View Synthesis

Examples of rendering new viewpoints

(5) Accessorization

Examples of adding accessories supported by the prior

(6) Property Modification

Examples of modifying semantically complex subject attributes

 

5.2. Ablation Studies

(1) Class-Prior Ablation

An incorrect class noun leaves an incompatible prior entangled with the subject, preventing effective new-image generation. Without a class noun, learning the instance and connecting it to the class prior becomes difficult, convergence takes longer, and errors increase.

(2) Prior Preservation Loss Ablation

The following shows prior-preservation loss preventing overfitting.

The following shows it preserving class-semantic priors.

(3) Super Resolution with Low-Noise Forward Diffusion

Lower noise augmentation improves sample quality and subject fidelity, as shown below.

5.3. Comparisons

Compared with An Image Is Worth One Word, this approach generates semantically accurate images and preserves subject features better.

 

5.4. Limitations

Several limitations remain, including three principal failure modes shown below.

  1. Generating images that do not match the prompt's context.
  1. Context-appearance entanglement
  1. Overfitting to prompts similar to the reference subject images.

Some subjects learn faster than others: common subjects have strong priors, while rare or complex subjects take longer. Fidelity also varies; depending on prior strength and the complexity of semantic changes, generated images can contain hallucinated subject features.

 

Reference

https://arxiv.org/abs/2112.10752

 

Read next