TEXTure: Text-Guided Texturing of 3D Shapes

TEXTure: Text-Guided Texturing of 3D Shapes

 
 

0. Abstact

TEXTure paints 3D objects using pretrained depth-to-image diffusion. Such models generate plausible textures from a single viewpoint, but stochastic generation causes inconsistency across the overall texturing process.

The paper divides rendered images into a trimap—generate, refine, and keep—and introduces sampling based on it. TEXTure can generate textures and also edit or refine existing ones through prompts or scribbles.

 

1. Introduction

The focus is seamlessly painting a supplied 3D mesh. Unlike earlier approaches, it directly applies a full denoising process to rendered images through depth-conditioned diffusion. It repeatedly renders different viewpoints, paints using depth, and projects results back onto mesh vertices or a texture atlas. This improves speed and quality, but naive application produces substantial inconsistency.

Dynamic partitioning converts each rendered view into a trimap of 'keep,' 'refine,' and 'generate.' Freezing keep regions improves consistency, but new regions still lack global consistency, as in Figure 2(B). Depth- and mask-guided diffusion improves generation, Figure 2(C), while a new process repaints refine regions. Together, these techniques produce highly realistic results in a few minutes.

The method accepts existing textures as well as text, without surface-to-surface mapping or explicit reconstruction. It extends An Image Is Worth One Word and DreamBooth to depth-conditioned models, learning texture-semantic tokens alongside viewpoint tokens. Text-driven and existing-map editing are both possible.

 

2. Related Work

  • Text-to-Image Diffusion Models
    • Introduces related techniques from Stable Diffusion and LDM through DreamBooth.

  • Texture and Content Transfer
    • Texturing a 3D surface is harder than 2D because both color and geometry matter. Earlier research covers geometric texture synthesis and 3D color-texture synthesis.

  • 3D Shape and Texture Generation
    • 3D shape and texture generation has gained attention. Text2Mesh, Tango, and CLIP-Mesh optimize CLIP-space similarity to generate shapes and textures. DreamFusion uses pretrained image diffusion to generate text-conditioned NeRFs. Its key is score-distillation loss, which optimizes 3D scenes with 2D priors. Latent-NeRF similarly applies score distillation in Stable Diffusion's latent space.

      For texture-only generation, Latent-Paint paints latent texture maps through score distillation and decodes them into RGB. Magic3D uses it to refine both texture and initial shape. This paper notes slow convergence and less-defined textures in both methods.

 

3. Method

3.1. Text-Guided Texture Synthesis

Texture generation uses depth-to-image model Mdepth\mathcal{M}_{depth} and inpainting model Mpaint\mathcal{M}_{paint}, both pretrained Stable Diffusion models sharing latent space. UV mapping represents the texture as an atlas, computed with XAtlas.

Start at arbitrary viewpoint v0=(r=1.25,ϕ0=0,θ=60)v_0 = (r=1.25, \phi_0=0, \theta=60), with camera radius rr, azimuth ϕ\phi, and elevation θ\theta. Model Mdepth\mathcal{M}_{depth} generates initial colored image I0I_0 from viewpoint v0v_0, conditioned on depth D0\mathcal{D}_0. Image I0I_0 projects into atlas T0\mathcal{T}_0. After initialization, incremental colorization follows fixed viewpoints, as in Figure 3.

At each viewpoint, renderer R\mathcal{R} produces Dt\mathcal{D}_t and QtQ_t. Rendering QtQ_t from viewpoint vtv_t incorporates all earlier colorization. Considering QtQ_t, the next image ItI_t and updated atlas Tt\mathcal{T}_t can be generated.

Once one view is painted, generation becomes harder because both local and global consistency matter. Let us examine one incremental painting iteration t and how it addresses this.

 

Trimap Creation

At viewpoint vtv_t, partition the rendered image into generate, keep, and refine regions. The generate region is newly visible and must be painted consistently with existing areas.

Distinguishing keep from refine is subtler. Oblique coloring angles can cause distortion through a small screen-space cross-section, producing low-resolution updates in texture image Tt\mathcal{T}_{t}. The triangle's cross-section is measured using component zz of its camera-space face normal nzn_z.

If the current view offers a better painting angle, refine the existing texture. Otherwise, keep it for consistency. An additional meta-texture mapN\mathcal{N} tracks previous regions and cross-sections, updating each iteration. It renders efficiently with the texture and determines the trimap.

 

Masked Generation

Depth-to-image diffusion generates entire images, so sampling must be modified to freeze keep regions. Following Blended Diffusion, noise the existing keep-region QtQ_t into zQtz_{Q_t} at each denoising step and inject it into sampling, blending seamlessly with generated output. The latent at step i is:

zi←zi⊙mblended+zQt⊙(1−mblended)z_i ← z_i \odot m_{blended} + z_{Q_{t}} \odot (1-m_{blended})

Mask mblendedm_{blended} is defined in Equation 2. For keep regions, simplyziz_i fix the original values.

 

Consistent Texture Generation

Injecting keep regions improves blending, but farther inside the generate region, sampled noise dominates and output can become inconsistent with earlier painting.

Inpainting model Mpaint\mathcal{M}_{paint}, trained to complete masked regions, improves consistency. Used alone, however, it can depart from depth Dt\mathcal{D}_t and invent geometry. An interleaved process combines both models, branching during initial sampling. The next noisy latent zi−1z_{i-1} is computed as follows.

Mdepth\mathcal{M}_{depth} guides noisy latents using current depth Dt\mathcal{D}_t, while Mpaint\mathcal{M}_{paint} completes generate regions in a globally consistent way.

 

Refining Regions

For refinement, diffusion must retain earlier values while adding texture. A checkerboard-like mask in early sampling guides noise toward those values. It applies during the first 25 steps, as follows.

A value of 1 means repaint the region; other values mean keep it. Figure 3 shows the blending mask.

 

Texture Projection

To project ItI_t into atlas Tt\mathcal{T}_t, apply gradient-based optimization to Lt\mathcal{L}_t with respect to Tt\mathcal{T}_t.

Soft mask msm_s smooths projection between views at refine/generate boundaries. gg is a 2D Gaussian blur kernel.

 

Additional Details

  • Textures use a 1024×1024 atlas.
  • Rendering resolution is 1200×1200.
  • For diffusion, resize the inner region to 512×512 and place it against a realistic background.
  • Render every shape from eight viewpoints plus top and bottom views.
 

3.2. Texture Transfer

Having generated new textures, the paper next transfers supplied textures onto unpainted meshes, either capturing them from painted meshes or from a few images. This builds on concept learning in An Image Is Worth One Word and fine-tuning in DreamBooth.

 

Spectral Augmentations

The goal is a token representing texture, not geometry, across shapes containing that texture. This requires disentangling the two and improving generalization. Spectral augmentation randomly inflates or contracts the mesh, as shown above.

 

Texture Learning

Spectral augmentation yields many image–depth pairs. Render several directions and composite onto colored backgrounds. Following An Image Is Worth One Word, optimize prompts of the form 'a <DvD_v> photo of a <StextureS_{texture}>.' Here, DvD_v represents view direction and StextureS_{texture} texture. Six DvD_v tokens are shared by matching directions, while StextureS_{texture} is shared across all images. DreamBooth fine-tuning captures texture better. After training, replace TEXTure's Stable Diffusion with the fine-tuned model to paint the target shape.

 

Texture from Images

Unlike standard textual inversion or DreamBooth, the learned concept primarily represents texture rather than structure learned by depth-conditioned diffusion. This may make it more suitable for 3D texturing.

For image-based transfer, pretrained U²-Net segments salient objects; scaling and cropping augment them before compositing onto randomly colored backgrounds. Semantic concepts learned from images transfer to 3D without explicit reconstruction, opening possibilities for textures inspired by real objects.

 

3.3. Texture-Editing

Text-based Editing

Trimap-based texturing extends 2D editing across the full mesh. For text-driven editing, mark the entire texture as refine, then align it with the new prompt through texturing.

 

Scribble-based Editing

Scribble-based editing lets users directly alter texture maps. Mark changed areas as refine and preserved areas as keep.

 

4. Experiments

4.1. Text-Guided Texturing

Qualitative Results

Qualitative Comparison

 

4.2. Texture Capturing

Texture From Mesh

 

Texture From Image

 

4.3. Editing

 

5. Limitation

  • Global inconsistency sometimes remains.
  • Eight viewpoints do not sufficiently cover challenging geometry.
    • Selecting enough viewpoints to maximize coverage could address this.
  • The depth-guided model sometimes departs from input depth and generates inconsistent images, shown on the left.
 

Reference

https://arxiv.org/abs/2302.01721

Read next