Prompt-to-Prompt Image Editing with Cross-Attention Control
Prompt-to-Prompt Image Editing with Cross-Attention Control
0. Abstract
Text-based diffusion poses a challenge for editing: even a small prompt change can completely alter the output. Existing state-of-the-art editing methods use masks, but do not preserve structure and content inside them.
The paper proposes prompt-to-prompt editing, controlling relationships between spatial layout and prompt words through cross-attention. It injects original attention maps into diffusion to control the edited image's maps.
Synthesis controlled entirely through textual prompts opens the way to varied caption-based editing applications.
1. Introduction
Large-scale language–image models have attracted researchers and the public with strong generation performance, but remain weak at controlling particular semantic regions in a given image.
Existing methods use masks for inpainting. Creating them is cumbersome and interrupts fast, intuitive text-driven editing. Removing important structural information also limits edits beyond inpainting, such as changing an object's texture.
Prompt-to-Prompt offers intuitive, powerful textual editing by manipulating internal cross-attention maps: high-dimensional tensors linking pixels with prompt tokens.
Injecting these maps into diffusion controls which pixels relate to which tokens at each step.

The paper proposes several methods for controlling cross-attention maps through a simple semantic interface; see Figure 1.
- Change one token, such as 'dog' to 'cat,' while freezing attention maps to preserve composition.
- Add words to introduce new attention flows while freezing attention for existing tokens.
- Strengthen or weaken a word's semantic effect on the image.
2. Related Work
GAN editing worked mainly on curated, narrow domains such as faces, not large diverse datasets. VQ-GAN enabled more diverse generation, while diffusion generally surpassed GAN performance. As diffusion improved, editing methods such as Blended Latent Diffusion emerged, but still required masks.
Unlike earlier work, this method needs only textual input, providing a more intuitive editing experience.
3. Method
Suppose image was generated from prompt and random seed . The goal is to edit with modified prompt to obtain . For an image generated from 'my new bicycle,' a user might change its color or replace it with a scooter while preserving structure. Editing the prompt is an intuitive interface.
To avoid masks, the authors first tried the simplest approach: use the same seed with the modified prompt. It failed, as Figure 2's bottom row shows.

The central discovery is that structure and appearance depend not only on the seed but also on pixel–text interactions during diffusion. Modifying those interactions in cross-attention enables Prompt-to-Prompt editing.
More precisely, injecting image 's cross-attention maps into diffusion preserves its composition and structure.
3.1. Cross-Attention in Text-Conditioned Diffusion Model
Imagen is the backbone, but the method is not model-specific. Section 4.1 also shows latent diffusion and Stable Diffusion results. All three condition on text through cross-attention. Since composition and geometry form at 6464 resolution, the method modifies only base diffusion and leaves super-resolution unchanged.

At diffusion step , a U-Net predicts noise from noisy image and text embedding , eventually producing image . Crucially, language and vision interact during noise prediction through cross-attention, which creates spatial attention maps for each token.
In Figure 3, deep spatial features of noisy image feed query matrix , while text embeddings feed key matrix and value matrix .
Attention maps are therefore:
Cell gives the weight of token for pixel , and is the latent projection dimension of keys and queries. Cross-attention output updates spatial features .
Intuitively, output is a weighted average of with weights . Parallel multi-head attention increases expressiveness.
3.2. Controlling the Cross-Attention

Returning to the discovery: spatial layout and geometry depend on cross-attention maps. Figure 4 shows pixels strongly associated with words describing them, such as 'bear.' Averaging is only for visualization; individual maps remain separate for their respective roles. Interestingly, visual structure is determined early in diffusion.
Because attention reflects composition, maps from original prompt can be injected into a second generation using edited prompt . Image then follows not only , but retains the structure of input . This is one form of attention manipulation; the paper proposes a more general framework.
Let denote one diffusion step , outputting noisy image and attention map . denotes a step overriding with while retaining supplementary prompt values . Maps for edited prompt are denoted , and the general editing function is .
The algorithm uses original and modified prompts simultaneously, fixing internal randomness because different seeds can produce entirely different outputs even for the same prompt. The algorithm follows.

Local Editing
Users generally want to edit one object or region while preserving other details, such as the background. The object's attention maps approximate an editing mask; see lines 11–14. At step , average using for original word and for new word . All experiments use .

Word Swap
For word replacement, users swap existing tokens. For example, = 'a big bicycle' becomes = 'a big car.' Preserving composition while adopting new content is difficult. The earlier injection overly constrains geometry, motivating the softer attention constraint below.

Timestamp parameter determines how long injection applies. Composition forms early, so limiting injection steps gives the new prompt more geometric freedom.
Prompt Refinement
Another case adds tokens. For example, = 'a castle' becomes = 'children drawing of a castle.' Injection applies only to shared tokens to preserve general details. Alignment function A maps a token index in target prompt to its matching index in , or None.

corresponds to pixels and to text tokens. Adjusting injection duration supports stylization, emphasizing object properties, and global manipulation.

Attetion Re-weighting
Finally, users may strengthen or weaken individual words' effects. For = 'a fluffy ball,' they might want more or less fluffiness. Parameter scales the attention map for token , leaving the others unchanged.

3.3. Self-Attention
Self-attention also affects layout and geometry, but involves only pixel-to-pixel interactions, unlike cross-attention. It cannot directly target edits associated with a particular text token.
4. Results
4.1. Applications

4.2. Comparisons

5. Limitations

- Current inversion visibly distorts some test images, as in Figure 19. It also requires a suitable prompt, which is difficult for complex compositions.
- Attention maps are low-resolution because cross-attention sits in the network bottleneck, limiting precise editing. Adding it to higher-resolution layers is proposed for future work.
- The method cannot move objects within an image.
6. Conclusions
The work reveals the power of diffusion's cross-attention layers. Their high-dimensional representations yield interpretable spatial maps connecting prompt words with generated layout.
This could be a first step toward simple, intuitive AI image editing.
Reference
https://arxiv.org/abs/2208.01626