arXiv:2512.22819cs.CV2025-12被引 1

让全景深度估计在真实世界中更准,无需重新训练。

Depth Anything in $360^\circ$: Towards Scale Invariance in the Wild

  • 从ViT骨干网络学习偏移参数,实现尺度不变的深度预测。
  • 引入环形填充消除接缝伪影,生成连续的全景深度图。
  • 在室内外数据集上均显著优于现有方法,适合机器人与虚实融合应用。

全景深度估计可全面捕捉360°环境结构信息,在机器人与AR/VR领域具有重要意义。然而,尽管室内场景研究充分,其在开放世界中的零样本泛化能力远落后于透视图像,后者得益于大量训练数据。为此,我们提出深度任意性全景版(DA360),基于Depth Anything V2进行全景适配。核心创新在于从ViT骨干网络学习一个偏移参数,将模型原本的尺度与平移不变输出转化为尺度不变估计,直接生成高质量3D点云。同时,在DPT解码器中引入环形填充,消除接缝伪影,确保空间一致性并满足球面连续性。在标准室内基准与新构建的室外数据集Metropolis上评估,相较于基线模型,DA360在室内和室外基准上分别实现超过50%和10%的相对误差降低。此外,相比现有稳健的全景深度估计方法,其在三个测试数据集上相对误差改善约30%,确立了零样本全景深度估计的新基准。

原文摘要 · Abstract (English)

Panoramic depth estimation provides a comprehensive solution for capturing complete $360^\circ$ environmental structural information, offering significant benefits for robotics and AR/VR applications. However, while extensively studied in indoor settings, its zero-shot generalization to open-world domains lags far behind perspective images, which benefit from abundant training data. This disparity makes transferring capabilities from the perspective domain an attractive solution. To bridge this gap, we present Depth Anything in $360^\circ$ (DA360), a panoramic-adapted version of Depth Anything V2. Our key innovation involves learning a shift parameter from the ViT backbone, transforming the model's scale- and shift-invariant output into a scale-invariant estimate that directly yields well-formed 3D point clouds. This is complemented by integrating circular padding into the DPT decoder to eliminate seam artifacts, ensuring spatially coherent depth maps that respect spherical continuity. Evaluated on standard indoor benchmarks and our newly curated outdoor dataset, Metropolis, DA360 shows substantial gains over its base model, achieving over 50\% and 10\% relative depth error reduction on indoor and outdoor benchmarks, respectively. Furthermore, DA360 significantly outperforms robust panoramic depth estimation methods, achieving about 30\% relative error improvement compared to PanDA across all three test datasets and establishing new state-of-the-art performance for zero-shot panoramic depth estimation.

全景深度零样本3D重建视觉感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。