An Image Is Worth One Word: Personalizing Text-to-Image Generation Using Textual Inversion
An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion
0. Abstract
How can a language-guided model turn the cat in a photograph into a painting, or design a new product based on our favorite toy?
This paper proposes learning new 'words' in embedding space to represent concepts such as objects or styles, using just three to five images supplied by the user. In particular, the authors find evidence that a single word embedding can capture unique and varied concepts.
1. Introduction


The method addresses earlier challenges by finding a single word in the textual embedding space of a pretrained text-to-image model. Consider the first stage of text encoding in Figure 2: an input string becomes tokens, each token is replaced by an embedding vector, and these vectors are fed into the downstream model. The goal is to find a new embedding vector for a particular concept.
The new embedding vector is represented by a new pseudo-word, denoted . It is treated like any other word: 'a photograph of on the beach.' We can also combine two concepts: 'a drawing of in the style of .'
To find this pseudo-word, the authors formulate an inversion task. Given a pretrained text-to-image model and a small set of three to five images, the goal is to find a single word embedding for 'A photo of .' Optimization finds this embedding. The process is called textual inversion.
To summarize, the contributions are:
- Introducing the task of personalized text-to-image generation.
- Proposing textual inversion to find new pseudo-words in the text encoder's embedding space.
- Analyzing the embedding space with GAN-inspired inversion techniques, identifying a trade-off between distortion and editability, and proposing an optimal point on that trade-off curve.
- Evaluating the method on images generated from user-provided concept captions, showing that the learned embeddings support higher visual fidelity and more robust editing.
3. Method
The goal is language-guided generation of new user-specified concepts. In conventional text-to-image models, candidate representations are found at the text encoder's word-embedding stage, but this approach does not require in-depth visual understanding of the images. The paper introduces a visual reconstruction objective inspired by GAN inversion.
Latent Diffusion Models
The method is applied to a latent diffusion model, using BERT as the text encoder .

Text embeddings
A typical text encoder such as BERT tokenizes text and embeds tokens as unique vectors, as shown below. The method chooses this embedding space as the inversion target and replaces the vector associated with S with v.

Textual inversion
Using just three to five images, it finds by minimizing the LDM loss in Equation 1. Generation is conditioned on context text randomly sampled from CLIP's ImageNet templates, such as 'A photo of ' and 'A rendition of .' The optimization goal is as follows.

To reuse the trained LDM, and are fixed—the diffusion model stays frozen—and only the text-encoder embedding is trained.
4. Qualitative comparison and applications
4.1. Image variations
In CLIP-based reconstruction, the method captures unique details and concepts well.

4.2. Text-guided synthesis

4.3. Style Transfer
Style transfer is also possible by replacing the context text with .

4.4. Concept Compositions
It can handle multiple words at once, but struggles to infer relationships between newly learned words.

4.5. Bias reduction
As Figure 8 shows, existing methods reinforce biases encoded in words such as 'doctor,' with most generated doctors being male. Learning a new embedding from a small but more diverse set can reduce those biases and improve representation of gender and racial diversity.

4.6. Downstream applications

6. Limitations
- The method offers greater freedom, but struggles to capture a concept's semantic essence or learn its exact shape.
- Optimization is slow: learning one concept takes about two hours.
Review
To me, its impact seems a little weaker than DreamBooth or Prompt-to-Prompt. Leaving the diffusion architecture untouched may be both an advantage and a limitation.
Reference
https://arxiv.org/abs/2208.01618