arXiv:2509.23991cs.CV2025-09NeurIPS被引 4

用全景图+图优化,零训练实现高精度360度深度估计

RPG360: Robust 360 Depth Estimation with Perspective Foundation Models and Graph Optimization

  • 将全景图转为六面体立方图,用视角大模型预测深度与法向
  • 通过图优化统一六面体间深度尺度,保持三维结构一致
  • 无需训练,在多个数据集上表现优越,适合下游视觉任务

随着360图像在各领域的广泛应用,针对全景图像的鲁棒深度估计需求日益迫切。然而,获取大规模带标注的360深度数据集仍是一大挑战。本文提出RPG360,一种无需训练的鲁棒单目360度深度估计方法,利用视角基础模型与图优化技术。方法将360图像转换为六面体立方图表示,使用视角基础模型估计深度和表面法向。为解决立方图不同面间深度尺度不一致问题,引入基于图优化的深度尺度对齐技术,参数化预测的深度与法向图,并引入每面独立的尺度参数。该优化确保六面体间深度尺度一致,同时保留三维结构完整性。由于基础模型在零样本场景下具备天然鲁棒性,本方法在Matterport3D、Stanford2D3D和360Loc等多个数据集上均表现优异。此外,验证了其在下游任务中的泛化优势:特征匹配提升3.2~5.4%,结构从运动(SfM)的AUC@5提升0.2~9.7%。

原文摘要 · Abstract (English)

The increasing use of 360 images across various domains has emphasized the need for robust depth estimation techniques tailored for omnidirectional images. However, obtaining large-scale labeled datasets for 360 depth estimation remains a significant challenge. In this paper, we propose RPG360, a training-free robust 360 monocular depth estimation method that leverages perspective foundation models and graph optimization. Our approach converts 360 images into six-face cubemap representations, where a perspective foundation model is employed to estimate depth and surface normals. To address depth scale inconsistencies across different faces of the cubemap, we introduce a novel depth scale alignment technique using graph-based optimization, which parameterizes the predicted depth and normal maps while incorporating an additional per-face scale parameter. This optimization ensures depth scale consistency across the six-face cubemap while preserving 3D structural integrity. Furthermore, as foundation models exhibit inherent robustness in zero-shot settings, our method achieves superior performance across diverse datasets, including Matterport3D, Stanford2D3D, and 360Loc. We also demonstrate the versatility of our depth estimation approach by validating its benefits in downstream tasks such as feature matching 3.2 ~ 5.4% and Structure from Motion 0.2 ~ 9.7% in AUC@5.

360度深度图优化基础模型全景图像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。