CraftsMan: High-fidelity Mesh Generation with 3D Native Generation and Interactive Geometry Refiner
Paper: https://arxiv.org/pdf/2405.14979 Project Page: https://craftsman3d.github.io/ Github: https://github.com/wyysf-98/CraftsMan
My thoughts on the paper

- Craftsmanship: Creativity and ingenuity
- Generates 3D meshes or geometry from image/text input through a two-stage, coarse-to-refine method comprising a 3D diffusion model and a geometry refiner.
- Coarse Stage
- The key idea is conditioning 3D-native diffusion with multi-view images.
- Naturally, these condition geometry better than a single image or a text prompt.
- It uses 2D techniques to complement the limitations of 3D-native methods: limited diversity and insufficient datasets.
- A well-trained multi-view diffusion model could be useful for both geometry and material generation.
- Refine Stage
- The refinement concept uses detail-enhanced normal maps for mesh optimization.
- Specifically, it fine-tunes a 2D diffusion model to generate normal images, then adds Tile ControlNet to refine their details. It optimizes the mesh using an L1 loss between the coarse mesh's normals rendered by a differentiable renderer and the normals refined by diffusion.
- Only the normal-generating diffusion component is trained.
- It is interesting how this seems to reverse traditional 3D modeling.
- In professional modeling, artists usually sculpt a high-poly mesh in ZBrush, retopologize it into a low-poly mesh, then bake normals with tools such as Substance Painter. Baking automatically extracts the difference between the two meshes into a normal map. Here, since generated geometry lacks detail, diffusion improves the normals and then transfers that detail back onto the mesh.
- If offered as a service, perhaps it could simply provide the enhanced normal map. Just a stray thought...
- Overall, the quality looks good. With CLAY's code still unavailable, this could be useful for research on the geometry-generation stage.
- For geometry alone, 3D-native methods with additional refinements might increasingly displace slow SDS-based approaches. The topology looks much better too.
1. Introduction

Creating 3D assets costs substantial time and money, so many approaches leverage generative methods to create 3D from just one image or a text prompt. These fall into three categories:
- 1. Score distillation sampling (SDS) methods
- Disadvantages: Slow generation, unstable optimization, and multiple faces—the Janus problem.
- 2. Multi-view (MV) methods
- Disadvantages: Irregular geometry, noisy surfaces, and over-smoothed surfaces.
- 3. 3D-native generation methods
- Disadvantages: Limited data and consequently weak generalization.
CraftsMan is a new 3D modeling system offering diverse shapes, regular mesh topology, detailed surfaces, and interactive geometry refinement from images or text. It has two stages: a coarse stage using a multi-view-conditioned 3D-native diffusion model, and a refinement stage using a generative geometry refiner. The native model learns the distribution of 3D geometry, but limited datasets inevitably hurt generalization. Combining it with multi-view diffusion addresses this. The refiner combines ControlNet-Tile with surface-normal diffusion to enhance details.
3. Method

The framework imitates an artist's workflow: build coarse geometry, then refine it, much like sculpting a high-poly model in ZBrush or Blender. This staged approach produces regular topology and detailed shapes. First, 3D assets are encoded into latents for shape-latent diffusion. Multi-view diffusion converts an input image or text into multiple views, which condition the 3D diffusion model. Surface-normal-based refinement comes last.
3.1. 3D Latent Set Representation
First, Figure (a) explains how the Shape VAE encodes 3D assets into latents.

Shape Representation
The success of latent diffusion models demonstrates the need for compact, efficient, expressive representations. Similar to LASER and PointFlow, the authors encode shapes as one-dimensional latent sets , where is the number of latent sets and the feature dimension.
Shape Encoding
An autoencoder encodes shapes into latent sets and reconstructs them as neural fields. For each shape, it samples a point cloud and normals from the surface. Cross-attention feeds Fourier positional encodings concatenated with normals into the shape encoder. A Perceiver-based encoder captures geometric properties effectively. ( supplies the two cross-attention inputs; denotes column-wise Fourier positional encoding.)

Shape Decoding
A similar Perceiver-based decoder moves all self-attention layers before cross-attention. Given a query point and learned shape latent embedding , it predicts an occupancy value, one type of 3D representation.

is the ground-truth value, while is an occupancy predictor consisting of a single MLP. , a KL-divergence loss, provides regularization. Marching Cubes reconstructs the final surface.
- Occupancy value
- An occupancy value indicates whether an object occupies a particular point in 3D space. It is usually binary, 0 or 1, or probabilistic, between 0 and 1.

3.2. 3D Native Diffusion Model
Uses multi-view images instead of directly conditioning on a single image or text prompt.

Generating MV Images as Intermediate Conditions
Multi-view images naturally provide richer geometric and contextual priors. They also unify the approach, avoiding separate training for text-conditioned and single-image-conditioned models. Looking at the code, it appears to support CRM, ImageDream, and Wonder3D implementations of multi-view diffusion, themselves drawing on approaches such as MVDream.

MV-conditioned 3D Latent Set Diffusion
The 3D diffusion model trains on shape latents and multi-view images . An image feature extractor embeds the views. Following Instant3D, camera parameters are included so the extractor can distinguish views. Applying adaLN to each attention sublayer makes the image embeddings camera-aware.

The 3D-native diffusion model thus generates an occupancy field conditioned on text or images, from which Marching Cubes extracts a mesh. To better display smooth shapes, a remeshing tool by Maxime [2024] converts triangular meshes into quadrilateral ones. Quad remeshing really seems to improve the quality.
3.3. Normal-based Geometry Refinement
Refinement strengthens coarse-mesh detail using normal maps. Rather than directly generating vertex changes, enhanced normal maps guide mesh optimization.

Coarse Normal Enhancement
ControlNet-Tile, commonly used for upscaling, improves rendered normal detail. The authors fine-tune a 2D diffusion model on normal images and use it with an RGB-trained ControlNet-Tile. Only normal-image generation is trained; the remaining components are reused.
Shape Optimization via Differentiable Rendering

The method uses vertex optimization with continuous remeshing from Palfinger, offering efficiency and explicit control. Given vertices , faces , and refined normal maps , it optimizes detail by directly manipulating triangle vertices and edges. Each step differentiably renders the current mesh's normals, then minimizes their L1 difference from the diffusion-refined target normals.

is the differentiable renderer, and is the camera information for view . Each step updates vertex positions according to the loss and splits, merges, or flips edges. See Palfinger for details.
Automatic Mesh Refinement
This multi-view refinement method is called Auto Normal Brush. Diffusion-generated normal images can be inconsistent across views, so it uses cross-view attention as in MVDream and Wonder3D. This issue is familiar when projecting multi-view images into textures: front and back views may work, while side views often fail to align. Linking attention keys and values transfers information between views and captures their relationships.
Retraining on synthetic datasets can overfit and produce unrealistic, over-smoothed results. Interestingly, the authors find that cross-view attention can be applied without training, thanks to coarse-normal constraints and ControlNet-Tile's ability to add detail without drastically changing the input. Refined normals then guide mesh optimization. I should test whether cross-view attention really works training-free.
Interactive Local Refinement
A method for refining only a user-selected region. Please refer to the paper.
4. Experiment
4.1. Dataset Preparation
- Filters low-quality meshes from 800,000 Objaverse assets, retaining 170,000.
- Each object is normalized to fit inside a unit cube, then preprocessed into a watertight mesh, as in Mesheder.
- Renders four orthogonal views under random rotations for the latent-set diffusion model.
- Uses the same approach for fine-tuning the 2D normal diffusion model.



5. Conclusion and Discussion
The original post labels this passage as a ChatGPT translation of the paper's overview: CraftsMan imitates a craftsperson's workflow and generates high-quality meshes in 30 seconds. It starts with coarse geometry and then enhances surface detail. A diffusion model trained directly on 3D geometry is conditioned on images produced by a strong multi-view diffusion model. This addresses scarce 3D data and improves robustness and generalization. A 2D diffusion model then enhances rendered normal maps, which guide detailed surface refinement.
Although the method generates high-quality meshes with regular topology, many directions remain open. Controllability of the latent-set diffusion model needs further study, and generating textures for 3D meshes is a promising topic for future work.