无需训练即可实现全景深度估计,保持几何一致性。
VGGT-360: Geometry-Consistent Zero-Shot Panoramic Depth Estimation
- 通过多视角重建3D模型,统一全景重投影推理。
- 在多个数据集上超越已有方法,支持多分辨率输入。
- 适合无训练场景下的全景深度生成任务。
本文提出VGGT-360,一种无需训练的零样本全景深度估计框架。与以往视图无关的免训练方法不同,VGGT-360利用VGGT类基础模型内在的3D一致性,将任务重构为基于多视角重建3D模型的全景重投影,从而将分散的单视角推理整合为连贯的全景理解。为实现鲁棒准确的估计,VGGT-360集成三个即插即用模块:(i) 基于不确定性的自适应投影,将全景图切分为透视视图,通过梯度不确定性分配更密集的视图至几何信息贫乏区域,提升输入质量;(ii) 结构显著性增强注意力,在3D重建中注入结构感知置信度,引导关注几何可靠区域,增强跨视图一致性;(iii) 相关性加权3D模型修正,利用注意力推断的相关性分数重加权重叠点,提供一致几何基础以支持精准全景重投影。大量实验表明,VGGT-360在多种分辨率和多样化的室内外数据集上,均优于现有训练及免训练的最先进方法。
原文摘要 · Abstract (English)
This paper presents VGGT-360, a novel training-free framework for zero-shot, geometry-consistent panoramic depth estimation. Unlike prior view-independent training-free approaches, VGGT-360 reformulates the task as panoramic reprojection over multi-view reconstructed 3D models by leveraging the intrinsic 3D consistency of VGGT-like foundation models, thereby unifying fragmented per-view reasoning into a coherent panoramic understanding. To achieve robust and accurate estimation, VGGT-360 integrates three plug-and-play modules that form a unified panorama-to-3D-to-depth framework: (i) Uncertainty-guided adaptive projection slices panoramas into perspective views to bridge the domain gap between panoramic inputs and VGGT's perspective prior. It estimates gradient-based uncertainty to allocate denser views to geometry-poor regions, yielding geometry-informative inputs for VGGT. (ii) Structure-saliency enhanced attention strengthens VGGT's robustness during 3D reconstruction by injecting structure-aware confidence into its attention layers, guiding focus toward geometrically reliable regions and enhancing cross-view coherence. (iii) Correlation-weighted 3D model correction refines the reconstructed 3D model by reweighting overlapping points using attention-inferred correlation scores, providing a consistent geometric basis for accurate panoramic reprojection. Extensive experiments show that VGGT-360 outperforms both trained and training-free state-of-the-art methods across multiple resolutions and diverse indoor and outdoor datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。