MVDream: Multi-View Diffusion for 3D Generation
MVDream: Multi-View Diffusion for 3D Generation
Paper: https://arxiv.org/abs/2308.16512 Project Page: https://dreamfusion3d.github.io/ Github: https://github.com/bytedance/MVDream

0. Abstract
This paper introduces MVDream, a multi-view diffusion model that generates consistent views from a text prompt. Training on both 2D and 3D data combines the generalizability of 2D diffusion with the consistency of 3D rendering. It also demonstrates improved representations when used as a prior for 3D generation. Through SDS, it improves the consistency and stability of 2D-lifting methods; see my DreamFusion review.
1. Introduction
3D content creation is essential to games and media, but is labor-intensive: a skilled designer spends substantial time on even one asset. A system enabling ordinary users without expertise in numerous tools to create 3D content easily would therefore be highly valuable.
Existing methods fall into (1) template-based pipelines and (2) 3D generative models, plus (3) 2D-lifting methods. Limited datasets and complex 3D data make the first two struggle with arbitrary objects, usually yielding simple topology and textures. Popular professional assets, however, are often complex, artistic, and sometimes unrealistic in structure and style.
Recent 2D-lifting methods, such as DreamFusion and Magic3D, optimize 3D representations through SDS using pretrained 2D models. Trained on large image datasets, these models generate novel or counterfactual scenes from text, making them excellent tools for artistic asset creation.

However, 2D-lifting methods lack comprehensive multi-view knowledge. This causes the multi-face Janus problem—repeated faces across views—and content drift between views. People can evaluate an object from multiple angles; conventional 2D diffusion cannot, leading to inconsistency.
Despite these limitations, the authors believe large 2D datasets can generalize to 3D. They propose multi-view diffusion as a representation-agnostic 3D prior that generates consistent views. Transfer learning inherits pretrained 2D diffusion's generalizability, then training combines rendered multi-view images with 2D image–text pairs. Used in SDS, it generates more stably than 2D diffusion while still creating unseen and counterfactual content. Inspired by DreamBooth, fine-tuning on a collection of 2D images remains robust. The model, MVDream, generates NeRFs without the usual multi-view issues and matches or exceeds previous state-of-the-art diversity.
2. Related Work and Background
2.1. 3D Generative Models
2.2. Diffusion Models for Object Novel View Synthesis
2.3. Lifting 2D Diffusion for 3D Generation
Given the bounded generalizability of 3D generative models, another thread of studies have attempted to apply 2D diffusion priors to 3D generation by coupling it with a 3D representation, such as a NeRF
3. Methodology
3.1. Multi-View Diffusion Model

A typical approach to inconsistency is improving viewpoint awareness. DreamFusion adds directional text such as 'back view.' Zero-1-to-3 more precisely conditions novel-view synthesis on camera parameters. The authors hypothesize that even perfect camera conditioning would not solve inconsistency: an eagle facing forward in a front view might still face right in a rear view.
Video diffusion provides another inspiration. People generally inspect multiple views to understand an object, much like a 360-degree turntable video. Recent video models generate temporally consistent output from 2D diffusion, but applying them to 3D does not necessarily ensure geometric consistency. The authors consider spatial consistency harder than temporal consistency. Video models also usually train on dynamic scenes, making static scenes difficult.
These observations lead to the conclusion that a multi-view diffusion model must be trained directly. A rendered 3D dataset can train camera-conditioned image generation.

Preserving a 2D diffusion architecture enables fine-tuning to inherit its flexibility. But it normally generates only one image and accepts no camera-position input. Three problems must be solved:
(1) How can one prompt generate consistent images from several angles? Section 3.1.1.
(2) How can camera-pose control be added? Section 3.1.2.
(3) How can quality and generalizability be preserved? Section 3.1.3.
3.1.1. Multi-View Consistent Image Generation
(a) As in video diffusion, attention could handle cross-view dependencies while retaining other networks. Simple temporal attention, however, failed to learn consistency; content drift remained even after fine-tuning on rendered 3D data.
(b) Adding a new 3D self-attention layer instead of modifying existing 2D attention also produced poor quality and weak multi-view generation. New modules converged more slowly. Since the authors wanted to avoid excessive fine-tuning, they abandoned this approach.
(c) Finally, they connected the original 2D self-attention layer across all views to make it 3D. They reused existing 2D attention for 3D attention.
3.1.2. Camera Embeddings
As in video diffusion, positional information distinguishes views. The authors compare relative position encoding, rotary embeddings, and absolute camera parameters. Camera parameters embedded with a two-layer MLP provide the best combination of view distinction and image quality. They test adding camera embeddings residually to time embeddings versus appending them to text embeddings for cross-attention. The former is more robust, apparently because it entangles camera and text less.
3.1.3. Training Loss Function
Data curation and training details are important; see the appendix. Stable Diffusion 2.1 is fine-tuned at 256×256. Combining a larger text-to-image dataset with rendered 3D data improves generalizability. The loss follows; see the paper for details.
3.2. Text-to-3D Generation
Multi-view diffusion can support 3D generation in two ways:
- Feed generated views into a few-shot reconstruction method, such as Instant-NGP.
- Use multi-view diffusion as the prior for SDS-based 2D lifting.
Reconstruction is more intuitive, but requires more views and stronger consistency, so the paper focuses on SDS. This requires changing camera sampling and accepting camera parameters as inputs. It uses the original text rather than DreamFusion-style direction-annotated prompts.
Although multi-view SDS generates consistent models, low-quality diffusion samples affect textures. The authors therefore anneal minimum and maximum timesteps, use fixed negative prompts, and reduce high-CFG color saturation with clamping techniques such as dynamic thresholding or CFG rescaling. They apply these tricks only to , retaining the SDS equation.
3.3. Multi-View DreamBooth for 3D Generation
See the paper.
4. Experiments
4.1. Multi-View Image Generation




4.2. 3D Generation with Multi-View Score Distillation

4.3. Multi-View DreamBooth
See the paper.
5. Discussion and Conclusion
Conclusion
The paper introduces a multi-view diffusion model. Training existing text-to-image diffusion on rendered 3D data and large-scale image–text data preserves generalizability while achieving multi-view consistency. It is also an effective SDS prior for 3D generation and performs well against other models.
Limitation
Training at 256×256 limits generalizability compared with the base model. An SDXL-based version is worth exploring. Lighting and textures also reflect the rendered dataset, so high-quality renderings are important.
Personal thoughts
- This feels like a major leap after the DreamFusion family of papers!
- I had been searching for ways to solve multi-view consistency for texture generation. Multi-view diffusion now offers a path.
- I should implement SDXL-based multi-view diffusion with ControlNet!