用谷歌地球合成模型对齐混乱的3D重建,解决无重叠视角下的全局一致性问题
Scene Grounding In the Wild
- 以伪合成场景为参考,通过语义高斯点云实现跨域对齐
- 在无视觉重叠条件下仍能保持全局一致,提升重建完整性
- 适用于大规模真实场景重建,尤其适合低重叠输入
从无序的野外图像中重建大规模真实场景的精确3D模型仍是计算机视觉的核心挑战,尤其是在输入视角之间几乎没有重叠的情况下。现有重建流程常产生多个不连通的局部重建,或错误地将非重叠区域合并为重叠几何结构。本文提出一种框架,将每个局部重建锚定到完整的参考模型上,实现在无视觉重叠条件下的全局一致对齐。参考模型来自由Google Earth Studio生成的密集、地理精准的伪合成渲染图,虽外观与真实照片差异显著,但共享相同的场景语义。我们采用3D高斯溅射表示参考模型,并为每个高斯点附加语义特征,将对齐建模为逆向特征优化问题,固定参考模型的同时估计全局6自由度位姿与尺度。此外,我们构建了WikiEarth数据集,将现有的部分3D重建与伪合成参考模型进行注册。实验表明,该方法在多种经典及基于学习的初始化方案下均能显著提升全局对齐效果,并缓解当前顶尖端到端模型的失效模式。
原文摘要 · Abstract (English)
Reconstructing accurate 3D models of large-scale real-world scenes from unstructured, in-the-wild imagery remains a core challenge in computer vision, especially when the input views have little or no overlap. In such cases, existing reconstruction pipelines often produce multiple disconnected partial reconstructions or erroneously merge non-overlapping regions into overlapping geometry. In this work, we propose a framework that grounds each partial reconstruction to a complete reference model of the scene, enabling globally consistent alignment even in the absence of visual overlap. We obtain reference models from dense, geospatially accurate pseudo-synthetic renderings derived from Google Earth Studio. These renderings provide full scene coverage but differ substantially in appearance from real-world photographs. Our key insight is that, despite this significant domain gap, both domains share the same underlying scene semantics. We represent the reference model using 3D Gaussian Splatting, augmenting each Gaussian with semantic features, and formulate alignment as an inverse feature-based optimization scheme that estimates a global 6DoF pose and scale while keeping the reference model fixed. Furthermore, we introduce the WikiEarth dataset, which registers existing partial 3D reconstructions with pseudo-synthetic reference models. We demonstrate that our approach consistently improves global alignment when initialized with various classical and learning-based pipelines, while mitigating failure modes of state-of-the-art end-to-end models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。