arXiv:2604.23432cs.CVcs.AI2026-04

针对全景相机姿态变化导致的深度估计失效问题,构建了首个系统性评测基准。

Sphere-Depth: A Benchmark for Depth Estimation Methods with Varying Spherical Camera Orientations

论文配图:Sphere-Depth: A Benchmark for Depth Estimation Methods with Varying Spherical Camera Orientations
图 1 · 摘自论文原文
  • 通过模拟全景相机姿态扰动,评估多模型在球面图像上的鲁棒性
  • 即使专为球面设计的模型,姿态偏移后性能仍显著下降
  • 提供校准协议与公开数据集,支持可复现评估

从全景图像中可靠地进行深度估计对机器人导航和沉浸式场景理解中的360°视觉至关重要。然而,在实际机器人平台中,机载全景相机可能产生非预期的姿态变化,加之等距柱状投影固有的几何失真,会严重影响深度估计效果。为此,本文提出一个名为Sphere-Depth的新公共基准,以可复现的方式系统评估单目深度估计模型在等距柱状图像上对相机姿态变化的鲁棒性。通过模拟相机姿态扰动,评估了主流透视基模型Depth Anything及球面感知模型Depth Anywhere、ACDNet、Bifuse++和SliceNet的表现。为进一步确保跨模型评估的合理性,提出基于深度校准的误差协议,利用每个模型的监督学习缩放因子将预测的相对深度转换为度量深度。实验表明,即便专为处理球面图像设计的模型,在相对于标准姿态出现偏差时仍表现出显著性能下降。完整基准、评估协议及数据集划分已公开于:https://github.com/sgazzeh/Sphere_depth

原文摘要 · Abstract (English)

Reliable depth estimation from spherical images is crucial for 360° vision in robotic navigation and immersive scene understanding. However, the onboard spherical camera can experience unintentional pose variations in real-world robotic platforms that, along with the geometric distortions inherent in equirectangular projections, significantly impact the effectiveness of depth estimation. To study this issue, a novel public benchmark, called Sphere-Depth, is introduced to systematically evaluate the robustness of monocular depth estimation models from equirectangular images in a reproducible way. Camera pose perturbations are simulated and used to assess the performance of a popular perspective-based model, Depth Anything, and of spherical-aware models such as Depth Anywhere, ACDNet, Bifuse++, and SliceNet. Furthermore, to ensure meaningful evaluation across models, a depth calibration-based error protocol is proposed to convert predicted relative depth values into metric depth values using supervised learned scaling factors for each model. Experiments show that even models explicitly designed to process spherical images exhibit substantial performance degradation when variations in the camera pose are observed with respect to the canonical pose. The full benchmark, evaluation protocol, and dataset splits are made publicly available at: https://github.com/sgazzeh/Sphere_depth

深度估计全景视觉机器人导航姿态鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。