用多视角信息增强扩散模型,让3D生成更可控。
DreamComposer++: Empowering Diffusion Models with Multi-View Conditions for 3D Content Generation
- 通过多视角特征融合提升3D表示能力
- 在多个视图条件下实现高可控性新视角合成
- 适配现有扩散模型,适合3D内容生成研究者
近期基于预训练2D扩散模型的方法已能从单张自然图像生成高质量新视角。然而,由于缺乏多视角信息,现有方法难以生成可控的新视角。本文提出DreamComposer++,一种灵活可扩展的框架,通过引入多视角条件来改进当前视图感知扩散模型。具体而言,该框架采用视图感知3D提升模块,从不同视角提取物体的3D表示,并通过多视角特征融合模块将这些表示聚合并渲染为目标视图的潜在特征。最终,目标视图特征被整合进预训练的图像或视频扩散模型中,用于新视角合成。实验表明,DreamComposer++可无缝集成至前沿视图感知扩散模型,显著提升其在多视角条件下的可控新视角生成能力,推动可控3D物体重建,适用于广泛应用场景。
原文摘要 · Abstract (English)
Recent advancements in leveraging pre-trained 2D diffusion models achieve the generation of high-quality novel views from a single in-the-wild image. However, existing works face challenges in producing controllable novel views due to the lack of information from multiple views. In this paper, we present DreamComposer++, a flexible and scalable framework designed to improve current view-aware diffusion models by incorporating multi-view conditions. Specifically, DreamComposer++ utilizes a view-aware 3D lifting module to extract 3D representations of an object from various views. These representations are then aggregated and rendered into the latent features of target view through the multi-view feature fusion module. Finally, the obtained features of target view are integrated into pre-trained image or video diffusion models for novel view synthesis. Experimental results demonstrate that DreamComposer++ seamlessly integrates with cutting-edge view-aware diffusion models and enhances their abilities to generate controllable novel views from multi-view conditions. This advancement facilitates controllable 3D object reconstruction and enables a wide range of applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。