用语义信息主动引导几何重建,提升稀疏视图下室内场景还原质量
AlignGS: Aligning Geometry and Semantics for Robust Indoor Reconstruction from Sparse Views
- 将语义先验作为几何正则化项,端到端联合优化几何与语义
- 在ScanNet等基准上新视图合成效果超越当前最优方法
- 适合需要高质量3D重建的AR/VR与机器人应用
构建语义丰富的室内3D模型需求日益增长,推动了增强现实、虚拟现实和机器人等领域的发展。然而,从稀疏视图中重建仍面临几何模糊问题。现有方法常将语义视为已生成几何上的被动标注,而本文提出语义应作为主动引导力量。为此,我们提出AlignGS框架,首次实现几何与语义的协同端到端优化。通过从2D基础模型中提取丰富先验,并设计深度一致性与多角度法向正则化等新型语义-几何引导机制,直接规约3D表示。在标准基准(如ScanNet)上的广泛评估表明,该方法在新视图合成任务中达到领先性能,且重建几何精度显著提升。结果验证了以语义先验作为几何正则化手段,可从有限输入视图生成更连贯、完整的3D模型。代码已开源:https://github.com/MediaX-SJTU/AlignGS。
原文摘要 · Abstract (English)
The demand for semantically rich 3D models of indoor scenes is rapidly growing, driven by applications in augmented reality, virtual reality, and robotics. However, creating them from sparse views remains a challenge due to geometric ambiguity. Existing methods often treat semantics as a passive feature painted on an already-formed, and potentially flawed, geometry. We posit that for robust sparse-view reconstruction, semantic understanding instead be an active, guiding force. This paper introduces AlignGS, a novel framework that actualizes this vision by pioneering a synergistic, end-to-end optimization of geometry and semantics. Our method distills rich priors from 2D foundation models and uses them to directly regularize the 3D representation through a set of novel semantic-to-geometry guidance mechanisms, including depth consistency and multi-faceted normal regularization. Extensive evaluations on standard benchmarks demonstrate that our approach achieves state-of-the-art results in novel view synthesis and produces reconstructions with superior geometric accuracy. The results validate that leveraging semantic priors as a geometric regularizer leads to more coherent and complete 3D models from limited input views. Our code is avaliable at https://github.com/MediaX-SJTU/AlignGS .
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。