arXiv:2505.10565cs.CV2025-05被引 40

融合不完整但精确与完整但相对的深度信息,生成任意场景的高精度稠密深度图。

Depth Anything with Any Prior

  • 分阶段融合度量深度与几何结构,通过像素级对齐和加权预填充先验
  • 在7个真实数据集上零样本表现优异,超越专用方法
  • 支持测试时切换模型,灵活调节精度与效率,适配新深度模型

本文提出Prior Depth Anything框架,将深度测量中不完整但精确的度量信息与深度预测中完整但相对的几何结构相结合,为任意场景生成高精度、稠密且细节丰富的度量深度图。为此,设计了从粗到精的渐进式融合流程:首先引入像素级度量对齐与距离感知加权,显式利用深度预测来预填充多样化的度量先验,有效缩小不同先验模式间的领域差异,提升跨场景泛化能力;其次,构建条件单目深度估计(MDE)模型,以归一化的预填充先验与预测作为条件,进一步隐式融合两类互补深度源,消除先验固有噪声。模型在7个真实世界数据集上实现出色的零样本泛化性能,覆盖深度补全、超分辨率与图像修复任务,效果媲美甚至超越此前专用方法。更重要的是,其在未见混合先验下仍表现稳健,并可通过更换预测模型实现测试时优化,提供灵活的精度-效率权衡,随MDE模型演进而持续升级。

原文摘要 · Abstract (English)

This work presents Prior Depth Anything, a framework that combines incomplete but precise metric information in depth measurement with relative but complete geometric structures in depth prediction, generating accurate, dense, and detailed metric depth maps for any scene. To this end, we design a coarse-to-fine pipeline to progressively integrate the two complementary depth sources. First, we introduce pixel-level metric alignment and distance-aware weighting to pre-fill diverse metric priors by explicitly using depth prediction. It effectively narrows the domain gap between prior patterns, enhancing generalization across varying scenarios. Second, we develop a conditioned monocular depth estimation (MDE) model to refine the inherent noise of depth priors. By conditioning on the normalized pre-filled prior and prediction, the model further implicitly merges the two complementary depth sources. Our model showcases impressive zero-shot generalization across depth completion, super-resolution, and inpainting over 7 real-world datasets, matching or even surpassing previous task-specific methods. More importantly, it performs well on challenging, unseen mixed priors and enables test-time improvements by switching prediction models, providing a flexible accuracy-efficiency trade-off while evolving with advancements in MDE models.

深度估计多源融合零样本通用性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。