DreamBooth: Fine-Tuning Text-to-Image Diffusion Models for Subject-Driven Generation
DreamBooth: Fine Tuning Text-to-Image Diffusion Models for Subject-Driven Generation
The GitHub link is a Stable Diffusion adaptation.

0. Abstact
Large text-to-image models generate high-quality images but cannot reliably reproduce subjects from a reference set. This paper fine-tunes them on a few subject images, binding a unique identifier to the subject. A class-specific prior-preservation loss leverages embedded semantic knowledge, enabling new scenes, poses, compositions, and lighting absent from the references.
1. Introduction
Synthesizing imaginary scenes has been challenging, but large text-to-image models have made remarkable progress. Their semantic priors link words such as 'dog' with many instances. Yet they cannot preserve a particular subject.
Their output domain is fundamentally limited. Even detailed descriptions produce different appearances, since text–object pairings are inherently diverse, as Figure 2 illustrates.

The paper proposes personalization by expanding the model's language–vision dictionary. Given three to five subject photos, its objective embeds that subject in the output domain, letting users generate it with a bound unique identifier.
Two new techniques are introduced:
- Represent a subject with a rare-token identifier.
- Fine-tune a diffusion-based text-to-image framework.
The framework has two stages—low-resolution generation followed by super-resolution—and fine-tuning likewise has two stages.
- Fine-tune low-resolution text-to-image generation with prompts containing a unique identifier, such as 'A [V] dog.'
- Apply class-specific prior-preservation loss to prevent overfitting and language drift.
- Fine-tune super-resolution.
- Preserve even small subject details with high fidelity.
The contributions are:
- Defining subject-driven generation: Given a few photos, generate the subject in varied contexts while preserving its key features with high fidelity.
- Proposing a few-shot fine-tuning technique for text-to-image diffusion that retains existing semantic knowledge.
2. Related Work
Related work considers whether subject-preserving novel generation could be solved without fine-tuning. Here is a brief summary of those alternatives.
(1) Compositing objects into scenes
Cloning a subject onto a new background cannot naturally adapt its lighting, interactions with objects, or shadows.
(2) Text-driven editing
CLIP-based GAN editing works on curated domains such as faces but struggles with diverse subject types. It handles global properties and local edits, not generating a specific subject in a new context.
(3) Text-to-image synthesis
Imagen, DALL·E 2, Parti, and CogView2 perform strongly but text alone offers limited fine control. Masks restrict areas, yet do not synthesize a subject in a new context.
(4) Diffusion-model inversion
This controls input latents, similar to manipulating W, W+, or S spaces in my GAN face-editing research. For diffusion, one must find a noise map or condition vector producing the desired image. It is a possible future solution, but currently limited by textual expressiveness and the pretrained model's output domain. See the papers below.
(5) Personalization: Fine-tuning for a specific domain
Personalization is common in recommendation and language models. GAN-based MyStyle fine-tunes for one facial identity, requiring around 100 photos. DreamBooth needs only three to five across varied subjects, including animals.
DreamBooth belongs to personalization!
3. Preliminaries
(1) Cascaded Text-to-Image Diffusion Models
Conditional diffusion model trains with squared error to denoise noisy image , as follows.
is ground truth, a condition vector from text or another source, and noise. , , and control the noise schedule and sample quality; time t follows . Inference repeatedly denoises using deterministic DDIM or stochastic ancestral sampling. predicts x.
State-of-the-art systems use cascaded diffusion: a 64×64 base text-to-image model followed by text-conditioned super-resolution from 6464 to 256256, then 256256 to 10241024.
(2) Vocabulary Encoding
Text conditioning is crucial for visual and semantic fidelity. Earlier models use CLIP text embeddings mapped into image embeddings by a learned prior, or pretrained T5-XXL.
This paper uses T5-XXL to embed tokenized prompts. Vocabulary encoding is important preprocessing. Prompt first passes through tokenizer before becoming condition embedding . The tokenizer is SentencePiece.
Tokenizer maps into fixed-length vector . Language model converts it to , and diffusion is conditioned on .
4. Method
The goal is new subject images with altered backgrounds, colors, species, shapes, poses, expressions, media, or other semantics. The high-level method follows.

First, embed the subject in the output domain and bind it to a unique identifier. Fine-tuning on a small set risks overfitting and language drift: forgetting other subjects in the same class and losing knowledge of its natural variations.
An autogenous class-specific prior-preservation loss reduces overfitting and drift by preserving generation of varied instances of the same class.
Super-resolution also needs fine-tuning to preserve detail. Naively training it to reproduce the instance misses important details. The paper provides training and testing insights for stronger preservation and recontextualization. Figure 4 shows the process; pretrained Imagen is the base.

4.1. Representing the Subject with a Rare-token Identifier
(1) Designing Prompts for Few-Shot Personalization
The goal is a new key–value pair in the dictionary, through few-shot fine-tuning. How should this be supervised? Human prompts and online data normally require detailed image descriptions, which are costly, variable, and subjective.
The paper instead labels images 'a [identifier] [class noun]'. The identifier is unique to the subject; the noun is a broad class, such as cat, dog, or watch, obtained from a classifier. Omitting it slows training and hurts performance. This labeling enables new poses and contexts.
Using existing words such as 'unique' or 'special' as identifiers requires disentangling their old meaning and attaching a new one, slowing learning and reducing performance. Using 'blue' for a gray subject, for example, mixes blue and gray in generated images.
An identifier with weak priors in both the language and diffusion models is needed. Random strings such as 'xxy5syt00' may appear literally in images or evoke related concepts, performing poorly like existing words.
(2) Rare-token Identifiers
The approach finds relatively rare vocabulary tokens and inverts them into text. Rare-token lookup gives identifier ; detokenizing yields text . Three or fewer Unicode characters without spaces work best. For example, , sks is vocabulary index 9,061 of 9,258.
4.2. Class-specific Prior Preserve Loss
(1) Few-shot Personalization of a Diffusion Model
Fine-tuning image set with common condition vector from 'a [identifier] [class noun]' and original loss causes overfitting and language drift.
(2) Issue-1: Overffiting
With few images, fine-tuning can overfit their appearance and context, as in Figure 12's top row. Regularization or selective layer tuning could help, but which layers should be frozen to balance fidelity and semantic flexibility is unclear. Empirically, tuning all layers works best but causes drift.

(3) Issue-2: Language Drift
Language-model research describes drift as gradually forgetting syntax and semantics during task-specific tuning. Here, after tuning on a few subject images, diffusion forgets the class prior and cannot generate other instances of that class, as in Figure 13's middle row.

(4) Prior-Preservation Loss
Prior-preservation loss supervises with samples generated by the original model, retaining its prior during few-shot tuning. A frozen model's ancestral sampler generates using initial noise and condition .
weights the preservation term. Figure 4 shows training. Despite its simplicity, the loss effectively addresses both problems. Learning rate and roughly epochs work best, taking around fifteen minutes on one TPU v4.
4.3. Personalized Instance-Specific Super-Resolution
The base model controls major visual semantics; super-resolution preserves details and photorealism. Without SR tuning, new subject images can contain artifacts, incorrect features, or lost detail, as in Figure 14's bottom row. Tuning 64×64 → 256×256 is essential; tuning 256×256 → 1024×1024 helps fine detail.

5. Experiments
5.1. Applications
(1) Recontextualization
Given model , prompts containing the identifier and class noun generate new subject images. Recontextualization usually uses 'a [V] [class noun] [context description]'.

(2) Art Renditions
Prompts such as 'a painting of a [V] [class noun] in the style of [famous painter]' or the equivalent for a sculpture produce artistic interpretations. This differs from transferring another image's style onto a fixed scene: it preserves identity and details while changing scene semantics, including poses and settings absent from the references.

(3) Expression Manipulation
Examples of generating new expressions

(4) Novel View Synthesis
Examples of rendering new viewpoints

(5) Accessorization
Examples of adding accessories supported by the prior

(6) Property Modification
Examples of modifying semantically complex subject attributes

5.2. Ablation Studies
(1) Class-Prior Ablation

An incorrect class noun leaves an incompatible prior entangled with the subject, preventing effective new-image generation. Without a class noun, learning the instance and connecting it to the class prior becomes difficult, convergence takes longer, and errors increase.
(2) Prior Preservation Loss Ablation
The following shows prior-preservation loss preventing overfitting.

The following shows it preserving class-semantic priors.

(3) Super Resolution with Low-Noise Forward Diffusion
Lower noise augmentation improves sample quality and subject fidelity, as shown below.

5.3. Comparisons
Compared with An Image Is Worth One Word, this approach generates semantically accurate images and preserves subject features better.

5.4. Limitations


Several limitations remain, including three principal failure modes shown below.
- Generating images that do not match the prompt's context.
- Context-appearance entanglement
- Overfitting to prompts similar to the reference subject images.
Some subjects learn faster than others: common subjects have strong priors, while rare or complex subjects take longer. Fidelity also varies; depending on prior strength and the complexity of semantic changes, generated images can contain hallucinated subject features.
Reference
https://arxiv.org/abs/2112.10752