Prompt-to-Prompt Image Editing with Cross-Attention Control

Prompt-to-Prompt Image Editing with Cross-Attention Control

 
 

0. Abstract

Text-based diffusion poses a challenge for editing: even a small prompt change can completely alter the output. Existing state-of-the-art editing methods use masks, but do not preserve structure and content inside them.

The paper proposes prompt-to-prompt editing, controlling relationships between spatial layout and prompt words through cross-attention. It injects original attention maps into diffusion to control the edited image's maps.

Synthesis controlled entirely through textual prompts opens the way to varied caption-based editing applications.

 

1. Introduction

Large-scale language–image models have attracted researchers and the public with strong generation performance, but remain weak at controlling particular semantic regions in a given image.

Existing methods use masks for inpainting. Creating them is cumbersome and interrupts fast, intuitive text-driven editing. Removing important structural information also limits edits beyond inpainting, such as changing an object's texture.

Prompt-to-Prompt offers intuitive, powerful textual editing by manipulating internal cross-attention maps: high-dimensional tensors linking pixels with prompt tokens.

Injecting these maps into diffusion controls which pixels relate to which tokens at each step.

The paper proposes several methods for controlling cross-attention maps through a simple semantic interface; see Figure 1.

  1. Change one token, such as 'dog' to 'cat,' while freezing attention maps to preserve composition.
  1. Add words to introduce new attention flows while freezing attention for existing tokens.
  1. Strengthen or weaken a word's semantic effect on the image.
 

2. Related Work

GAN editing worked mainly on curated, narrow domains such as faces, not large diverse datasets. VQ-GAN enabled more diverse generation, while diffusion generally surpassed GAN performance. As diffusion improved, editing methods such as Blended Latent Diffusion emerged, but still required masks.

Unlike earlier work, this method needs only textual input, providing a more intuitive editing experience.

 

3. Method

Suppose image I\mathcal{I} was generated from prompt P\mathcal{P} and random seed ss. The goal is to edit I\mathcal{I} with modified prompt P∗\mathcal{P}^* to obtain I∗\mathcal{I}^*. For an image generated from 'my new bicycle,' a user might change its color or replace it with a scooter while preserving structure. Editing the prompt is an intuitive interface.

To avoid masks, the authors first tried the simplest approach: use the same seed with the modified prompt. It failed, as Figure 2's bottom row shows.

The central discovery is that structure and appearance depend not only on the seed but also on pixel–text interactions during diffusion. Modifying those interactions in cross-attention enables Prompt-to-Prompt editing.

More precisely, injecting image I\mathcal{I}'s cross-attention maps into diffusion preserves its composition and structure.

 

3.1. Cross-Attention in Text-Conditioned Diffusion Model

Imagen is the backbone, but the method is not model-specific. Section 4.1 also shows latent diffusion and Stable Diffusion results. All three condition on text through cross-attention. Since composition and geometry form at 64×\times64 resolution, the method modifies only base diffusion and leaves super-resolution unchanged.

 

At diffusion step tt, a U-Net predicts noise ϵ\epsilon from noisy image ztz_t and text embedding ψ(P)\psi(\mathcal{P}), eventually producing image I=z0\mathcal{I} = z_0. Crucially, language and vision interact during noise prediction through cross-attention, which creates spatial attention maps for each token.

In Figure 3, deep spatial features of noisy image ϕ(zt)\phi(z_t) feed query matrix Q=lQ(ϕ(zt))Q=l_Q(\phi(z_t)), while text embeddings feed key matrix K=lK(ψ(P))K = l_K(\psi(P)) and value matrix V=lV(ψP)V=l_V(\psi{P}).

Attention maps are therefore:

M=softmax(QKTd)M = softmax(\cfrac {QK^T}{\sqrt{d}})

Cell MijM_{ij} gives the weight of token jj for pixel ii, and dd is the latent projection dimension of keys and queries. Cross-attention output ϕ^(zt)=MV\hat{\phi}(z_t) = MV updates spatial features ϕ(zt)\phi(z_t).

Intuitively, output MVMV is a weighted average of VV with weights MM. Parallel multi-head attention increases expressiveness.

 

3.2. Controlling the Cross-Attention

Returning to the discovery: spatial layout and geometry depend on cross-attention maps. Figure 4 shows pixels strongly associated with words describing them, such as 'bear.' Averaging is only for visualization; individual maps remain separate for their respective roles. Interestingly, visual structure is determined early in diffusion.

Because attention reflects composition, maps MM from original prompt P\mathcal{P} can be injected into a second generation using edited prompt P∗\mathcal{P}^*. Image I∗\mathcal{I}^* then follows not only P∗\mathcal{P}^*, but retains the structure of input I\mathcal{I}. This is one form of attention manipulation; the paper proposes a more general framework.

Let DM(zt,P,t,s)DM(z_t,\mathcal{P},t,s) denote one diffusion step tt, outputting noisy image zt−1z_{t-1} and attention map Mt M_t. DM(zt,P,t,s){M←M^}DM(z_t,\mathcal{P},t,s)\{ M \leftarrow \hat{M} \} denotes a step overriding MM with M^\hat{M} while retaining supplementary prompt values VV. Maps for edited prompt P∗\mathcal{P^*} are denoted Mt∗M_t^*, and the general editing function is Edit(Mt,Mt∗,t)Edit(M_t, M_t^*, t).

The algorithm uses original and modified prompts simultaneously, fixing internal randomness because different seeds can produce entirely different outputs even for the same prompt. The algorithm follows.

 

Local Editing

Users generally want to edit one object or region while preserving other details, such as the background. The object's attention maps approximate an editing mask; see lines 11–14. At step tt, average T,…,tT,…,t using Mˉt,w\bar{M}_{t,w} for original word ww and M∗ˉt,w∗\bar{M^*}_{t,w^*} for new word w∗w^*. All experiments use B(x):=x>k,  k=0.3B(x) := x> k, \ \ k=0.3.

 

Word Swap

For word replacement, users swap existing tokens. For example, P\mathcal{P} = 'a big bicycle' becomes P∗\mathcal{P}^* = 'a big car.' Preserving composition while adopting new content is difficult. The earlier injection overly constrains geometry, motivating the softer attention constraint below.

Timestamp parameter τ\tau determines how long injection applies. Composition forms early, so limiting injection steps gives the new prompt more geometric freedom.

 

Prompt Refinement

Another case adds tokens. For example, P\mathcal{P} = 'a castle' becomes P∗\mathcal{P}^* = 'children drawing of a castle.' Injection applies only to shared tokens to preserve general details. Alignment function A maps a token index in target prompt P∗\mathcal{P}^* to its matching index in P\mathcal{P}, or None.

ii corresponds to pixels and jj to text tokens. Adjusting injection duration supports stylization, emphasizing object properties, and global manipulation.

 

Attetion Re-weighting

Finally, users may strengthen or weaken individual words' effects. For P\mathcal{P} = 'a fluffy ball,' they might want more or less fluffiness. Parameter c∈[−2,2]c \in [-2,2] scales the attention map for token j∗j^*, leaving the others unchanged.

 

3.3. Self-Attention

Self-attention also affects layout and geometry, but involves only pixel-to-pixel interactions, unlike cross-attention. It cannot directly target edits associated with a particular text token.

 

4. Results

4.1. Applications

 

4.2. Comparisons

 
 

5. Limitations

  1. Current inversion visibly distorts some test images, as in Figure 19. It also requires a suitable prompt, which is difficult for complex compositions.
  1. Attention maps are low-resolution because cross-attention sits in the network bottleneck, limiting precise editing. Adding it to higher-resolution layers is proposed for future work.
  1. The method cannot move objects within an image.
 

6. Conclusions

The work reveals the power of diffusion's cross-attention layers. Their high-dimensional representations yield interpretable spatial maps connecting prompt words with generated layout.

This could be a first step toward simple, intuitive AI image editing.

 

Reference

https://arxiv.org/abs/2208.01626

 
 

Read next