用多视角引导和表面密化生成高质量3D模型,半小时完成训练。
MVGaussian: High-Fidelity text-to-3D Content Generation with Multi-View Guidance and Surface Densification
- 通过多视角引导逐步构建3D结构,提升细节与准确性。
- 提出新密化算法使高斯点贴近表面,增强模型保真度。
- 仅需30分钟训练即达高质结果,效率远超现有方法。
文本到3D内容生成领域已取得显著进展,现有方法如得分蒸馏采样(SDS)提供了有效指导。然而,这些方法常面临“双面人”问题——因指导不精确导致多面歧义。尽管近期3D高斯分裂在表示3D体积方面表现出色,但其优化仍缺乏深入探索。本文提出一个统一框架,解决上述关键缺陷。方法利用多视角引导迭代构建3D模型结构,逐步提升细节与精度;同时引入新型密化算法,将高斯点对齐至表面附近,优化模型结构完整性和保真度。大量实验验证了该方法的有效性,生成结果视觉质量高,训练耗时极低:仅需半小时即可达到与多数现有方法数小时训练相当的性能,显著提升效率。
原文摘要 · Abstract (English)
The field of text-to-3D content generation has made significant progress in generating realistic 3D objects, with existing methodologies like Score Distillation Sampling (SDS) offering promising guidance. However, these methods often encounter the "Janus" problem-multi-face ambiguities due to imprecise guidance. Additionally, while recent advancements in 3D gaussian splitting have shown its efficacy in representing 3D volumes, optimization of this representation remains largely unexplored. This paper introduces a unified framework for text-to-3D content generation that addresses these critical gaps. Our approach utilizes multi-view guidance to iteratively form the structure of the 3D model, progressively enhancing detail and accuracy. We also introduce a novel densification algorithm that aligns gaussians close to the surface, optimizing the structural integrity and fidelity of the generated models. Extensive experiments validate our approach, demonstrating that it produces high-quality visual outputs with minimal time cost. Notably, our method achieves high-quality results within half an hour of training, offering a substantial efficiency gain over most existing methods, which require hours of training time to achieve comparable results.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。