TEXTure: Text-Guided Texturing of 3D Shapes
TEXTure: Text-Guided Texturing of 3D Shapes

paper: https://arxiv.org/abs/2302.01721 project page: https://texturepaper.github.io/TEXTurePaper/
0. Abstact
TEXTure paints 3D objects using pretrained depth-to-image diffusion. Such models generate plausible textures from a single viewpoint, but stochastic generation causes inconsistency across the overall texturing process.
The paper divides rendered images into a trimap—generate, refine, and keep—and introduces sampling based on it. TEXTure can generate textures and also edit or refine existing ones through prompts or scribbles.
1. Introduction
The focus is seamlessly painting a supplied 3D mesh. Unlike earlier approaches, it directly applies a full denoising process to rendered images through depth-conditioned diffusion. It repeatedly renders different viewpoints, paints using depth, and projects results back onto mesh vertices or a texture atlas. This improves speed and quality, but naive application produces substantial inconsistency.

Dynamic partitioning converts each rendered view into a trimap of 'keep,' 'refine,' and 'generate.' Freezing keep regions improves consistency, but new regions still lack global consistency, as in Figure 2(B). Depth- and mask-guided diffusion improves generation, Figure 2(C), while a new process repaints refine regions. Together, these techniques produce highly realistic results in a few minutes.
The method accepts existing textures as well as text, without surface-to-surface mapping or explicit reconstruction. It extends An Image Is Worth One Word and DreamBooth to depth-conditioned models, learning texture-semantic tokens alongside viewpoint tokens. Text-driven and existing-map editing are both possible.
2. Related Work
- Text-to-Image Diffusion Models
Introduces related techniques from Stable Diffusion and LDM through DreamBooth.
- Texture and Content Transfer
Texturing a 3D surface is harder than 2D because both color and geometry matter. Earlier research covers geometric texture synthesis and 3D color-texture synthesis.
- 3D Shape and Texture Generation
3D shape and texture generation has gained attention. Text2Mesh, Tango, and CLIP-Mesh optimize CLIP-space similarity to generate shapes and textures. DreamFusion uses pretrained image diffusion to generate text-conditioned NeRFs. Its key is score-distillation loss, which optimizes 3D scenes with 2D priors. Latent-NeRF similarly applies score distillation in Stable Diffusion's latent space.
For texture-only generation, Latent-Paint paints latent texture maps through score distillation and decodes them into RGB. Magic3D uses it to refine both texture and initial shape. This paper notes slow convergence and less-defined textures in both methods.
3. Method

3.1. Text-Guided Texture Synthesis
Texture generation uses depth-to-image model and inpainting model , both pretrained Stable Diffusion models sharing latent space. UV mapping represents the texture as an atlas, computed with XAtlas.
Start at arbitrary viewpoint , with camera radius , azimuth , and elevation . Model generates initial colored image from viewpoint , conditioned on depth . Image projects into atlas . After initialization, incremental colorization follows fixed viewpoints, as in Figure 3.
At each viewpoint, renderer produces and . Rendering from viewpoint incorporates all earlier colorization. Considering , the next image and updated atlas can be generated.
Once one view is painted, generation becomes harder because both local and global consistency matter. Let us examine one incremental painting iteration t and how it addresses this.
Trimap Creation
At viewpoint , partition the rendered image into generate, keep, and refine regions. The generate region is newly visible and must be painted consistently with existing areas.
Distinguishing keep from refine is subtler. Oblique coloring angles can cause distortion through a small screen-space cross-section, producing low-resolution updates in texture image . The triangle's cross-section is measured using component of its camera-space face normal .
If the current view offers a better painting angle, refine the existing texture. Otherwise, keep it for consistency. An additional meta-texture map tracks previous regions and cross-sections, updating each iteration. It renders efficiently with the texture and determines the trimap.
Masked Generation
Depth-to-image diffusion generates entire images, so sampling must be modified to freeze keep regions. Following Blended Diffusion, noise the existing keep-region into at each denoising step and inject it into sampling, blending seamlessly with generated output. The latent at step i is:
Mask is defined in Equation 2. For keep regions, simply fix the original values.
Consistent Texture Generation
Injecting keep regions improves blending, but farther inside the generate region, sampled noise dominates and output can become inconsistent with earlier painting.
Inpainting model , trained to complete masked regions, improves consistency. Used alone, however, it can depart from depth and invent geometry. An interleaved process combines both models, branching during initial sampling. The next noisy latent is computed as follows.

guides noisy latents using current depth , while completes generate regions in a globally consistent way.
Refining Regions
For refinement, diffusion must retain earlier values while adding texture. A checkerboard-like mask in early sampling guides noise toward those values. It applies during the first 25 steps, as follows.

A value of 1 means repaint the region; other values mean keep it. Figure 3 shows the blending mask.
Texture Projection
To project into atlas , apply gradient-based optimization to with respect to .
Soft mask smooths projection between views at refine/generate boundaries. is a 2D Gaussian blur kernel.

Additional Details
- Textures use a 1024×1024 atlas.
- Rendering resolution is 1200×1200.
- For diffusion, resize the inner region to 512×512 and place it against a realistic background.
- Render every shape from eight viewpoints plus top and bottom views.
3.2. Texture Transfer
Having generated new textures, the paper next transfers supplied textures onto unpainted meshes, either capturing them from painted meshes or from a few images. This builds on concept learning in An Image Is Worth One Word and fine-tuning in DreamBooth.
Spectral Augmentations

The goal is a token representing texture, not geometry, across shapes containing that texture. This requires disentangling the two and improving generalization. Spectral augmentation randomly inflates or contracts the mesh, as shown above.
Texture Learning

Spectral augmentation yields many image–depth pairs. Render several directions and composite onto colored backgrounds. Following An Image Is Worth One Word, optimize prompts of the form 'a <> photo of a <>.' Here, represents view direction and texture. Six tokens are shared by matching directions, while is shared across all images. DreamBooth fine-tuning captures texture better. After training, replace TEXTure's Stable Diffusion with the fine-tuned model to paint the target shape.
Texture from Images
Unlike standard textual inversion or DreamBooth, the learned concept primarily represents texture rather than structure learned by depth-conditioned diffusion. This may make it more suitable for 3D texturing.
For image-based transfer, pretrained U²-Net segments salient objects; scaling and cropping augment them before compositing onto randomly colored backgrounds. Semantic concepts learned from images transfer to 3D without explicit reconstruction, opening possibilities for textures inspired by real objects.
3.3. Texture-Editing
Text-based Editing
Trimap-based texturing extends 2D editing across the full mesh. For text-driven editing, mark the entire texture as refine, then align it with the new prompt through texturing.
Scribble-based Editing
Scribble-based editing lets users directly alter texture maps. Mark changed areas as refine and preserved areas as keep.
4. Experiments
4.1. Text-Guided Texturing
Qualitative Results


Qualitative Comparison

4.2. Texture Capturing
Texture From Mesh

Texture From Image

4.3. Editing

5. Limitation

- Global inconsistency sometimes remains.
- Eight viewpoints do not sufficiently cover challenging geometry.
- Selecting enough viewpoints to maximize coverage could address this.
- The depth-guided model sometimes departs from input depth and generates inconsistent images, shown on the left.