DreamFusion: Text-to-3D Using 2D Diffusion

DreamFusion: Text-to-3D using 2D Diffusion

 
 

0. Abstract

Recent advances in text-to-image synthesis owe much to diffusion models. Applying them directly to 3D requires large labeled 3D datasets and architectures that efficiently denoise 3D data, neither of which is currently available.

This paper proposes a loss based on probability-density distillation that makes a 2D diffusion model usable for image-generation optimization. Gradient descent can optimize a randomly initialized 3D model, a NeRF, so that 2D renderings from random angles receive low loss. Neither 3D data nor modifications to the diffusion model are required.

 

1. Introduction

The paper introduces a technique that converts a text-to-image diffusion model into a guide for 3D synthesis without 3D data. Diffusion works in other modalities, but normally requires large modality-specific training datasets. Such a dataset is unavailable for 3D.

2D generation is now readily accessible, but games and films need highly detailed 3D assets, still painstakingly designed in Blender or Maya. Text-to-3D models could lower barriers for beginners and improve experienced artists' workflows.

 

Earlier 3D synthesis research

  • Voxels and point clouds: Learn explicit structural representations, but require 3D data.
  • GANs: Learn controllable 3D generators, but only within the specific trained domain.
  • NeRF: Performs classical 3D reconstruction. Given images of a scene, it recovers geometry and generates views from previously unseen angles.
  • NeRF-like models: Dream Fields uses CLIP's joint image–text embeddings and an optimization-based approach to train NeRF, showing that 2D image–text models can guide 3D synthesis.
  • DreamFusion: Similar to Dream Fields, but replaces CLIP with a loss distilled from 2D diffusion. Combining SDS with NeRF enables high-quality object generation from prompts alone.
 

2. Diffusion Models and Score Distillation Sampling

2.1. How can we sample in parameter space, not pixel space

Diffusion samplers normally generate samples of the same type and dimensionality as their training data. Conditional sampling adds flexibility such as inpainting, but pixel-trained models have still been used to sample pixels. This paper instead seeks a 3D model whose renderings from random angles look like good images.

This can be formulated through differentiable image parameterization (DIP): a differentiable generator gg maps parameters θ\theta to images x=g(θ)\mathrm{x}=g(\theta). DIP expresses constraints, enables optimization in a compact space, and leverages strong pixel-space optimization. For 3D,θ\theta parameterizes a volume, while gg is a volumetric renderer. Training these parameters requires a diffusion-based loss. Naturally, good images should receive low loss and poor ones high loss.

Initially, the authors tried the diffusion training loss, but it did not produce realistic samples. Some studies emphasized careful timestep schedules; here, the objective itself proved fragile and schedule tuning difficult. Examine the gradient of LDiff\mathcal{L}_{Diff} to understand why.

The U-Net Jacobian is expensive to compute and poorly conditioned at low noise levels. Omitting that term makes DIP optimization effective.

Intuitively, the loss perturbs image x with random noise and estimates an update direction following the diffusion model's score function. Although the gradient may appear ad hoc, it is the gradient of a probability-density distillation loss; see Appendix A.4. The authors call the approach score distillation sampling (SDS).

LSDS\mathcal{L}_{SDS} is easy to implement, as in Figure 8, and w(t)w(t) can be chosen relatively robustly. The model predicts the update direction directly, so no backpropagation through diffusion is needed. The mode-seeking behavior of LSDS\mathcal{L}_{SDS} makes good samples seem uncertain, but Figure 2 shows reasonable-quality SDS images.

 

3. The DreamFusion Algorithm

We have seen how diffusion supplies a loss for a general continuous optimization problem, LSDS\mathcal{L}_{SDS}. For text-to-3D, the paper uses the unchanged 64×64 base Imagen model, without super-resolution. A NeRF-like model starts with random weights. Repeated renderings from different camera positions and angles become inputs to LSDS\mathcal{L}_{SDS}, which wraps Imagen; see Figure 3.

3.1. Neural Rendering of a 3D Model

NeRF combines a volumetric ray tracer with an MLP for neural inverse rendering. Rays travel from the camera center through image pixels. Sampled 3D points μ\mu along each ray pass through the MLP, yielding four scalars:

  • τ\tau: Volumetric density, representing the opacity of scene geometry at a 3D position.
  • cc: RGB color, three scalar values.

Densities and colors are alpha-composited along the ray toward the camera to obtain the rendered RGB value.

Conventional NeRF receives image–camera pairs and trains a randomly initialized MLP using MSE between rendered and ground-truth pixel colors. The resulting 3D model, parameterized by MLP weights, renders realistic unseen views. DreamFusion builds on mip-NeRF 360, an improved NeRF that reduces aliasing. Though designed for reconstruction, it also benefits text-to-3D synthesis.

Shading

Conventional NeRF models emitted light, making RGB values dependent on ray direction. DreamFusion instead parameterizes intrinsic color and controls its shading. Earlier NeRF-like models used various reflectance models; this paper uses per-point RGB albedo ρ\rho, the material's color.

(τ,ρ)=MLP(μ;θ)(\tau, \rho) = MLP(\mu; \theta)

Computing shaded output requires a normal vector describing local surface orientation. It is obtained through n=−∇μτ∣∣∇μτ∣∣ n=-\nabla_\mu \tau ||\nabla_\mu \tau||. With normal nn, albedo ρ\rho, light position ll, light color lρl_\rho, and ambient color lal_a, we can calculate each point's color along the ray.

These colors and estimated densities approximate the volume-rendering integral. Following earlier work, randomly setting albedo ρ\rho to white (1, 1, 1) produces textureless shading. This discourages flat geometry that merely satisfies the prompt—for example, encouraging a 3D squirrel rather than a flat squirrel image with suitable angle and lighting.

 

Scene Structure

Although the approach can generate complex scenes, the authors find it helpful to query NeRF only inside a fixed bounding sphere and use an environment map generated by a second MLP from positionally encoded ray directions. Accumulated alpha blends ray color over the background. This colors the scene's surroundings without encouraging density immediately beside the camera. A smaller bounding sphere can help when generating a single object.

 

Geometry regularizers

mip-NeRF 360 includes many other details omitted here. Regularization discourages filling empty space. A modified Ref-NeRF orientation loss prevents undesirable density-field behavior. It matters particularly for textureless shading: without it, normals may turn away from the camera and make shading dark. See Appendix A.2.

 

3.2. Text-to-3D Synthesis

With a pretrained text-to-image model, a NeRF-based DIP, and the loss, we are ready for text-to-3D synthesis without 3D data. Each prompt trains a randomly initialized NeRF from scratch. Every iteration follows these steps; pseudocode appears in Appendix 8.

  • (1) randomly sample a camera and light
    • Randomly sample a camera and lighting.
  • (2) render an image of the NeRF from that camera and shade with the light
    • Render NeRF using the sampled camera and shading.
  • (3) compute gradients of the SDS loss with respect to the NeRF parameters
    • Calculate the SDS loss.
  • (4) update the NeRF parameters using an optimizer
    • Update NeRF parameters with an optimizer.
 

See the paper for details.

4. Experiments

  • Comparison with other zero-shot text-to-3D models
  • Evaluation uses CLIP R-Precision, measuring consistency between rendered images and the input caption.
 

Ablations

 

5. Discussion

DreamFusion synthesizes 3D from varied text prompts. SDS and a new NeRF-like rendering engine turn a 2D diffusion model into a 3D generation guide, requiring neither 3D nor multi-view training data.

 

Limitations

  • SDS is not a perfect image-sampling loss: compared with ancestral sampling, it often produces oversaturated, over-smoothed results. Dynamic thresholding partially mitigates this for images, but not for NeRF. SDS images also lack diversity, and DreamFusion's 3D results differ little across random seeds. Reverse KL divergence and its mode-seeking behavior may explain this.
  • Using 64×64 Imagen limits detail. Higher-resolution diffusion and larger NeRFs might help, but would make synthesis impractically slow. More efficient diffusion and neural rendering could enable practical high-resolution 3D synthesis.
  • Reconstructing 3D from 2D observations has no unique solution. This ambiguity affects generation. Fundamentally, the task is difficult for the same reasons inverse rendering is difficult.
 

Read next