用视频扩散模型生成3D物体,让单图转3D更一致、更逼真。
NOVA3D: Normal Aligned Video Diffusion Model for Single Image to 3D Generation

- 用视频扩散模型提供强3D先验,提升多视角一致性。
- 提出GTA注意力机制,实现颜色与几何信息高效融合。
- 适合想快速生成高质量3D模型的创作者和研究人员。
3D AI生成内容(AIGC)使普通人也能成为3D内容创作者。尽管近期方法利用得分蒸馏采样从预训练图像扩散模型中提炼3D物体,但常因3D先验不足导致多视角一致性差。本文提出NOVA3D,一种创新的单图生成3D框架。核心思路是利用预训练视频扩散模型中的强3D先验,并在多视角视频微调中融入几何信息。为促进颜色与几何域间的信息交换,提出几何-时间对齐(GTA)注意力机制,提升泛化性与多视角一致性。此外,引入去冲突几何融合算法,通过解决多视角误差和姿态对齐差异,提高纹理保真度。大量实验验证了NOVA3D在性能上优于现有基线方法。
原文摘要 · Abstract (English)
3D AI-generated content (AIGC) has made it increasingly accessible for anyone to become a 3D content creator. While recent methods leverage Score Distillation Sampling to distill 3D objects from pretrained image diffusion models, they often suffer from inadequate 3D priors, leading to insufficient multi-view consistency. In this work, we introduce NOVA3D, an innovative single-image-to-3D generation framework. Our key insight lies in leveraging strong 3D priors from a pretrained video diffusion model and integrating geometric information during multi-view video fine-tuning. To facilitate information exchange between color and geometric domains, we propose the Geometry-Temporal Alignment (GTA) attention mechanism, thereby improving generalization and multi-view consistency. Moreover, we introduce the de-conflict geometry fusion algorithm, which improves texture fidelity by addressing multi-view inaccuracies and resolving discrepancies in pose alignment. Extensive experiments validate the superiority of NOVA3D over existing baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。