CLAY: A Controllable Large-scale Generative Model for Creating High-quality 3D Assets
CLAY: A Controllable Large-scale Generative Model for Creating High-quality 3D Assets
Paper: https://arxiv.org/abs/2406.13897 Project Page: https://sites.google.com/view/clay-3dlm Github: https://github.com/CLAY-3D/OpenCLAY Service: https://hyperhuman.deemos.com/rodin (Rodin)
CLAY: A Controllable Large-scale Generative Model for Creating High-quality 3D Assets My thoughts on the paper 1. Introduction3. Large-Scale 3D Generative Model3.1. Representation and Model Architecture3.2. Data Standardization for Pretraining4. Asset Enhancement5. Model Adaptation6. Result

My thoughts on the paper
- Geometry Generative Model
- Establishes a new baseline for geometry generation.
- Improves performance with a scalable diffusion-transformer architecture.
- Selects 527,000 high-quality 3D assets from Objaverse and ShapeNet, then remeshes and labels them.
- Tests various model sizes, with mesh topology clearly better than previous methods. 👍
- Material Diffusion Model
- Selects 40,000 good PBR assets from Objaverse.
- Modifies MVDream, a multi-view diffusion model, to train four-view material diffusion.
- Adds normal ControlNet and LoRA, but does not explain this part in detail...
- This seems to preserve texture-to-mesh alignment well, but needs verification.
- Generates diffuse, roughness, and metallic maps using TEXTure's gradient-based texture-projection optimization.
- When I tried the service, color was cleanly separated into the diffuse map. 👍
- Why did they not call it albedo?
- Fills occluded regions using Text2Tex's automatic viewpoint selection.
- The service shows almost no occlusion gaps. Is this really how they address them? Needs verification.
- Personal impressions
- It seems worthy of a SIGGRAPH Best Paper award: novelty, performance, and a good system all together.
- I liked the two-stage concept of geometry generation plus material generation. It is surprising to finally see plausible PBR results.
- But is 3D generation ultimately becoming a contest between scalable transformer-based models and large datasets? If so, I am personally wondering how to approach future research.
- Will it come down to capital and data, as with LLMs?
- Or will a new approach prevail because 3D and graphics lack sufficient real-world training data?
- I will need to keep watching the direction of the field.
1. Introduction
- There are two broad strategies for 3D generation.
- [A] Extending 2D generative models into 3D, for example through SDS
- Advantage: Diverse generation results.
- Disadvantage: Limited understanding of 3D, making high geometric fidelity difficult to maintain.
- [B] Learning from 3D datasets
- Advantage: Better understanding and preservation of geometric fidelity.
- Disadvantage: Large datasets are difficult to obtain, which is closely related to the problem 3D generation itself seeks to solve.
- This paper combines [A], 2D-based generation, and [B], 3D-based generation, in a pretrain-then-adapt paradigm.
- Geometry generation
- 3D native geometry generator(1.5B)
- Components: 3DShape2VecSet, a VAE, and a diffusion transformer (DiT).
- Uses a new 3D preprocessing pipeline: remeshing and GPT-4V labeling.
- Material generation
- Generates four views with a multi-view material diffusion model, then projects them.
- Trained on high-quality PBR textures(diffuse, roughness, and metallic) from Objaverse
3. Large-Scale 3D Generative Model
- Challenges in developing a 3D generative model
- How should a 3D model be represented?
- Type-1. geometry with per-vertex color
- Type-2. geometry with texture map
- Should geometry be extracted from 3D appearance representations such as NeRF or Gaussian splatting, or generated directly?
- CLAY simply separates geometry generation from texture generation: Type 2.
- This means not using 2D generation to create 3D geometry.
- Experiments show that scaling the 3D geometry model and its dataset can outperform earlier 2D-based models in both diversity and quality.
- Simply put, CLAY is a large 3D generative model with 1.5 billion parameters.
- Architecture-wise
- CLAY extends the generative model of 3DShape2VecSet with a new multi-resolution VAE.
- This makes geometry encoding and decoding much more efficient.
- It also adds an improved diffusion transformer for probabilistic geometry generation.
- Dataset-wise
- Develops a remeshing pipeline.
- Develops a GPT-4V annotation schema to standardize and unify existing 3D datasets.
- Different formats and inconsistent data had previously limited their use.
3.1. Representation and Model Architecture

- The approach to 3D generation
- Like 2D generative models, it focuses on learning to denoise data in a compressed latent space.
- Effectively reduces the complexity and computational cost of 3D space.
- Uses 3DShape2VecSet's representation and architecture as a basis, adapted for scaling up.
- 1. Encode 3D geometry into latent space
- Sample a point cloud from the mesh surface .
- Encode the point cloud into a latent code with a dynamic shape of × 64.
- The encoder is a transformer-based VAE.
- 2. Denoise with DiT
- noise at step t
- 3. Decode into a neural field through the VAE decoder
- Isosurface
- Multi-resolution VAE
- See the paper.
- Model scalability is important, making the multi-resolution VAE's support for dynamic shapes essential.
- Coarse-to-fine DiT
- See the paper for experiments with various model sizes.
- This progressive scaling method ensures robust and efficient training of our DiT
- Training details are provided.
- Scaling-up Scheme
- See the paper.
- Scaling-up CLAY requires enhancing both the VAE and DiT architectures with pre-normalization and GeLU activation, to facilitate faster computation of attention mechanism.
- Our largest model, the XL, was trained on a cluster of 256 NVidia A800 GPUs, for approximately 15 days, with progressive training
- Following the insights in Gesmundo and Maile [2023] of Head addition, Heads expansion and Hidden dimension expansion, we progressively scale up the DiT during training. This approach offers benefits such as enhanced time efficiency, improved knowledge retention, and a reduced risk of the model trapped in the local optima
- That is what the paper reports.

3.2. Data Standardization for Pretraining
- The effectiveness and robustness of large-scale 3D generative models depend on dataset quality and scale. Unlike text or images, open 3D datasets are scarce.
- Obtaining large quantities of high-quality 3D data requires addressing non-watertight meshes, inconsistent orientations, and inaccurate annotations.
- CLAY's solution is a remeshing method combined with GPT-4V annotation.
- Standardization
- Filtering unsuitable data
- Unsuitable data includes complex scenes and fragmented scans.
- Filtering yields 527,000 assets from ShapeNet and Objaverse.
- Geometry Unification
- Non-watertight, or open, meshes remain even after filtering.
- The authors therefore experiment with remeshing using Manifold, mesh-to-sdf, and DOGN—the Dual Octree Graph Network.
- Remeshing requires the following strategies:
- (1) Geometric preservation: Maintain important geometric features with minimal alteration.
- (2) Volume preservation: Ensure the integrity of all structural components.
- (3) Adaptability to non-watertight meshes.
- Inspired by DOGN, they adopt an unsigned distance field (UDF) representation.
- An additional issue occurs when extracting an isosurface through Marching Cubes from a mesh with holes: only thin surfaces are extracted.
- we employ a grid-based visibility computation before isosurface extraction S
- Geometry Annotation
- As Stable Diffusion demonstrates, prompts are important.
- The authors develop their own prompt tags and generate detailed annotations using GPT-4V.
- This improves the model's ability to interpret and generate complex 3D geometry with fine detail and diverse styles.
4. Asset Enhancement
- A two-stage method makes results directly usable in a CG pipeline.
- post-generation geometry optimization + material synthesis
- Geometry optimization refines the mesh.
- Material synthesis creates realistic textures.
- Mesh Quadrification and Atlasing
- CLAY's geometry extracted through Marching Cubes consists of uneven triangles.
- Blender converts triangular faces into quadrilateral faces.
- Quadrification is important for obtaining the final mesh.
- Material Synthesis
- PBR materials consist of diffuse, metallic, and roughness maps.
- Personally, unlike CLAY, I think two maps—albedo and ORM (occlusion, roughness, metallic)—might work, if possible...
- Earlier PBR generation methods handle only a small subset of these materials.
- RichDreamer generates only diffuse maps, without roughness or metallic maps.
- And so on.
- Carefully selects 40,000 objects from Objaverse.
- Unfortunately, there is no detailed explanation of how they selected them or which metadata they used.
- These objects have high-quality PBR materials.
- Uses this dataset to develop multi-view material diffusion.
- Much faster than previous methods.
- Maps results into UV space similarly to TEXTure—through UV optimization?
- Modifies MVDream to develop four-view material diffusion.
- This accommodates additional channels and modalities.
- Modifies the U-Net, inspired by HyperHuman's facial texture generator.
- See my DreamFace review!
- Trains on orthogonal views and adds LoRA fine-tuning.
- Uses normal ControlNet, which seems important for recognizing geometric detail.
- Besides preserving geometric accuracy, this also supports image inputs through mechanisms such as IP-Adapter.
- After generation with four-view material diffusion:
- Inpaints using adaptive view selection, as in Text2Tex.
- It does not inpaint in UV space. Hmm...
- Applies super-resolution to upscale to 2K.
- Uses Real-ESRGAN and MultiDiffusion techniques.
- A question: How did they achieve mesh alignment?

5. Model Adaptation
- Explains how various controls are injected.
- Pass
6. Result
- Essentially, the results say it is the best.


