Text2Tex: Text-driven Texture Synthesis via Diffusion Models
Text2Tex: Text-driven Texture Synthesis via Diffusion Models
paper: https://arxiv.org/abs/2303.11396 project page: https://daveredrum.github.io/Text2Tex/
0. Abstract
This paper introduces Text2Tex, which uses a depth-aware image diffusion model to generate high-quality textures for a given 3D mesh.
1. Introduction
Despite recent successes in generating 3D geometry, the need for manually designed textures still prevents fully automated 3D content creation. We therefore need to automate texture design using alternative guidance, such as text.
Text2Tex uses a generate-then-refine strategy. It progressively generates partial textures from different viewpoints and back-projects them into texture space. To address inconsistent artifacts visible from other viewpoints—for example, defects revealed when rotating away from a generated front view—it introduces a view partitioning technique. This calculates similarity maps between texel normals and the current viewing direction.
It also introduces an automatic viewpoint selection technique to ensure full coverage of the mesh surface while producing high-quality, consistent texture maps. The technique progressively selects the best view for the next step.
3. Method
3.2. Depth-Aware Image Inpainting
The core of texture synthesis is painting unfilled areas of the mesh surface. For inpainting, this paper uses a generation mask , similar to TEXTure (see Masked Generation in Section 3.1). Refer to TEXTure for details.
3.3. Progressive Texture Generation
To synthesize a texture for the input mesh, UV parameterization projects the 2D view—the image generated by Stable Diffusion—into the texture space of the normalized 3D object.
A viewpoint is defined by , where is the azimuth angle about the axis, is the elevation angle relative to the plane, and is the distance between the object and the origin.
This is also similar to TEXTure, apart from the notation. See Section 3.1., Text-Guided Texture Synthesis and Trimap Creation.
3.4. Texture Refinement with Automatic Viewpoint Selection

4. Result
4.1. Imprementation Details
- Uses the Depth2Image model from Stable Diffusion v2.
- Denoising strengths and are set to 0.5 for generation and 0.3 for refinement, respectively.
- Uses six axis-aligned viewpoints for generation.
- For refinement, selects only 20 of 36 viewpoints to reduce runtime.
- Each synthesis process takes 15 minutes.
4.2. Experiments Setup
- Evaluation uses the Objaverse dataset.
4.3. Quantitivative results


4.4. Qualitative analysis


4.5. Ablation studies
Does depth-aware inpainting and updating help?
- yes


Does viewpoint selection in refinement stage help?
- yes

4.6. Limitations
The proposed method can generate high-quality 3D textures, but the authors observed shading effects introduced by the diffusion backbone. These can be controlled through input prompts, but doing so requires manual engineering. One possible solution is to fine-tune the diffusion model to avoid generating shading in textures.
Review
Perhaps because the papers appeared around the same time, this is remarkably similar to TEXTure—even the methods are almost identical. Since the code is not available either, it seems better to refer to TEXTure.
Reference
https://daveredrum.github.io/Text2Tex/static/Text2Tex.pdf