CLAY: A Controllable Large-scale Generative Model for Creating High-quality 3D Assets

CLAY: A Controllable Large-scale Generative Model for Creating High-quality 3D Assets

 
 
CLAY 모델을 기반으로한 Deemos Technology의 Text/Image to 3D, Rodin 서비스
Rodin, Deemos Technology's text/image-to-3D service based on CLAY

My thoughts on the paper

  • Geometry Generative Model
    • Establishes a new baseline for geometry generation.
    • Improves performance with a scalable diffusion-transformer architecture.
    • Selects 527,000 high-quality 3D assets from Objaverse and ShapeNet, then remeshes and labels them.
    • Tests various model sizes, with mesh topology clearly better than previous methods. 👍
  • Material Diffusion Model
    • Selects 40,000 good PBR assets from Objaverse.
    • Modifies MVDream, a multi-view diffusion model, to train four-view material diffusion.
      • Adds normal ControlNet and LoRA, but does not explain this part in detail...
      • This seems to preserve texture-to-mesh alignment well, but needs verification.
    • Generates diffuse, roughness, and metallic maps using TEXTure's gradient-based texture-projection optimization.
      • When I tried the service, color was cleanly separated into the diffuse map. 👍
      • Why did they not call it albedo?
    • Fills occluded regions using Text2Tex's automatic viewpoint selection.
      • The service shows almost no occlusion gaps. Is this really how they address them? Needs verification.
  • Personal impressions
    • It seems worthy of a SIGGRAPH Best Paper award: novelty, performance, and a good system all together.
    • I liked the two-stage concept of geometry generation plus material generation. It is surprising to finally see plausible PBR results.
    • But is 3D generation ultimately becoming a contest between scalable transformer-based models and large datasets? If so, I am personally wondering how to approach future research.
      • Will it come down to capital and data, as with LLMs?
      • Or will a new approach prevail because 3D and graphics lack sufficient real-world training data?
      • I will need to keep watching the direction of the field.
 

1. Introduction

  • There are two broad strategies for 3D generation.
    • [A] Extending 2D generative models into 3D, for example through SDS
      • Advantage: Diverse generation results.
      • Disadvantage: Limited understanding of 3D, making high geometric fidelity difficult to maintain.
    • [B] Learning from 3D datasets
      • Advantage: Better understanding and preservation of geometric fidelity.
      • Disadvantage: Large datasets are difficult to obtain, which is closely related to the problem 3D generation itself seeks to solve.
  • This paper combines [A], 2D-based generation, and [B], 3D-based generation, in a pretrain-then-adapt paradigm.
    • Geometry generation
      • 3D native geometry generator(1.5B)
        • Components: 3DShape2VecSet, a VAE, and a diffusion transformer (DiT).
      • Uses a new 3D preprocessing pipeline: remeshing and GPT-4V labeling.
    • Material generation
      • Generates four views with a multi-view material diffusion model, then projects them.
      • Trained on high-quality PBR textures(diffuse, roughness, and metallic) from Objaverse
       

3. Large-Scale 3D Generative Model

  • Challenges in developing a 3D generative model
    • How should a 3D model be represented?
      • Type-1. geometry with per-vertex color
      • Type-2. geometry with texture map
    • Should geometry be extracted from 3D appearance representations such as NeRF or Gaussian splatting, or generated directly?
  • CLAY simply separates geometry generation from texture generation: Type 2.
    • This means not using 2D generation to create 3D geometry.
    • Experiments show that scaling the 3D geometry model and its dataset can outperform earlier 2D-based models in both diversity and quality.
  • Simply put, CLAY is a large 3D generative model with 1.5 billion parameters.
    • Architecture-wise
      • CLAY extends the generative model of 3DShape2VecSet with a new multi-resolution VAE.
      • This makes geometry encoding and decoding much more efficient.
      • It also adds an improved diffusion transformer for probabilistic geometry generation.
    • Dataset-wise
      • Develops a remeshing pipeline.
      • Develops a GPT-4V annotation schema to standardize and unify existing 3D datasets.
        • Different formats and inconsistent data had previously limited their use.
 

3.1. Representation and Model Architecture

  • The approach to 3D generation
    • Like 2D generative models, it focuses on learning to denoise data in a compressed latent space.
    • Effectively reduces the complexity and computational cost of 3D space.
    • Uses 3DShape2VecSet's representation and architecture as a basis, adapted for scaling up.
      • 1. Encode 3D geometry into latent space
        • Sample a point cloud XX from the mesh surface MM .
        • Encode the point cloud XX into a latent code ZZ with a dynamic shape of LL × 64.
        • The encoder is a transformer-based VAE.
      • 2. Denoise ZZ with DiT
        • noise at step t
      • 3. Decode into a neural field through the VAE decoder
        • Isosurface
    • Multi-resolution VAE
      • See the paper.
      • Model scalability is important, making the multi-resolution VAE's support for dynamic shapes essential.
    • Coarse-to-fine DiT
      • See the paper for experiments with various model sizes.
      • This progressive scaling method ensures robust and efficient training of our DiT
      • Training details are provided.
    • Scaling-up Scheme
      • See the paper.
      • Scaling-up CLAY requires enhancing both the VAE and DiT architectures with pre-normalization and GeLU activation, to facilitate faster computation of attention mechanism.
      • Our largest model, the XL, was trained on a cluster of 256 NVidia A800 GPUs, for approximately 15 days, with progressive training
      • Following the insights in Gesmundo and Maile [2023] of Head addition, Heads expansion and Hidden dimension expansion, we progressively scale up the DiT during training. This approach offers benefits such as enhanced time efficiency, improved knowledge retention, and a reduced risk of the model trapped in the local optima
      • That is what the paper reports.
      •  

3.2. Data Standardization for Pretraining

  • The effectiveness and robustness of large-scale 3D generative models depend on dataset quality and scale. Unlike text or images, open 3D datasets are scarce.
  • Obtaining large quantities of high-quality 3D data requires addressing non-watertight meshes, inconsistent orientations, and inaccurate annotations.
  • CLAY's solution is a remeshing method combined with GPT-4V annotation.
  • Standardization
    • Filtering unsuitable data
      • Unsuitable data includes complex scenes and fragmented scans.
      • Filtering yields 527,000 assets from ShapeNet and Objaverse.
    • Geometry Unification
      • Non-watertight, or open, meshes remain even after filtering.
      • The authors therefore experiment with remeshing using Manifold, mesh-to-sdf, and DOGN—the Dual Octree Graph Network.
      • Remeshing requires the following strategies:
        • (1) Geometric preservation: Maintain important geometric features with minimal alteration.
        • (2) Volume preservation: Ensure the integrity of all structural components.
        • (3) Adaptability to non-watertight meshes.
      • Inspired by DOGN, they adopt an unsigned distance field (UDF) representation.
      • An additional issue occurs when extracting an isosurface through Marching Cubes from a mesh with holes: only thin surfaces are extracted.
        • we employ a grid-based visibility computation before isosurface extraction S
    • Geometry Annotation
      • As Stable Diffusion demonstrates, prompts are important.
      • The authors develop their own prompt tags and generate detailed annotations using GPT-4V.
      • This improves the model's ability to interpret and generate complex 3D geometry with fine detail and diverse styles.
 

4. Asset Enhancement

  • A two-stage method makes results directly usable in a CG pipeline.
  • post-generation geometry optimization + material synthesis
    • Geometry optimization refines the mesh.
    • Material synthesis creates realistic textures.
  • Mesh Quadrification and Atlasing
    • CLAY's geometry extracted through Marching Cubes consists of uneven triangles.
    • Blender converts triangular faces into quadrilateral faces.
    • Quadrification is important for obtaining the final mesh.
 
  • Material Synthesis
    • PBR materials consist of diffuse, metallic, and roughness maps.
      • Personally, unlike CLAY, I think two maps—albedo and ORM (occlusion, roughness, metallic)—might work, if possible...
    • Earlier PBR generation methods handle only a small subset of these materials.
      • RichDreamer generates only diffuse maps, without roughness or metallic maps.
      • And so on.
    • Carefully selects 40,000 objects from Objaverse.
      • Unfortunately, there is no detailed explanation of how they selected them or which metadata they used.
      • These objects have high-quality PBR materials.
      • Uses this dataset to develop multi-view material diffusion.
        • Much faster than previous methods.
      • Maps results into UV space similarly to TEXTure—through UV optimization?
    • Modifies MVDream to develop four-view material diffusion.
      • This accommodates additional channels and modalities.
      • Modifies the U-Net, inspired by HyperHuman's facial texture generator.
      • Trains on orthogonal views and adds LoRA fine-tuning.
      • Uses normal ControlNet, which seems important for recognizing geometric detail.
      • Besides preserving geometric accuracy, this also supports image inputs through mechanisms such as IP-Adapter.
      • After generation with four-view material diffusion:
        • Inpaints using adaptive view selection, as in Text2Tex.
          • It does not inpaint in UV space. Hmm...
        • Applies super-resolution to upscale to 2K.
          • Uses Real-ESRGAN and MultiDiffusion techniques.
        • A question: How did they achieve mesh alignment?
 

5. Model Adaptation

  • Explains how various controls are injected.
  • Pass
 

6. Result

  • Essentially, the results say it is the best.
 

Read next