通过自适应尺度场提升单目深度估计的度量精度。
Learning Image-Adaptive Scale Fields for Metric Depth Recovery

- 用图像自适应基图的线性组合建模尺度变化,替代直接修正深度。
- 在极稀疏度量锚点下仍保持高精度,误差显著降低。
- 方法可解释且适用于多种主流单目深度模型。
单目深度估计(MDE)通常仅能恢复相对深度,缺乏度量尺度。当仅有稀疏度量锚点时,恢复准确度量深度极具挑战但对实际应用至关重要。本文将度量深度恢复问题建模为图像自适应尺度场,不直接修正深度,而是将其表示为由语义与几何线索生成的图像自适应基图的低维线性组合。基图权重通过最小二乘法从稀疏度量锚点高效求解。该方法显著提升度量深度精度,在极端锚点稀疏条件下仍具强鲁棒性,并实现空间尺度变化的可解释分解。在多个数据集和代表性MDE模型上的大量实验验证了其有效性与通用性。
原文摘要 · Abstract (English)
Monocular depth estimation (MDE) typically produces depth estimations that are defined up to an unknown scale or shift. When only sparse metric anchors are available, recovering accurate metric depth becomes challenging yet necessary for practical applications. We address this problem by formulating metric depth recovery as image-adaptive scale field modeling. Instead of directly correcting the depth, we reformulate the correction as a low-dimensional linear combination of image-adaptive basis maps. These maps are derived from semantic and geometric cues encoded in the MDE estimations and intermediate representations. The weights of basis maps are efficiently determined from sparse metric anchors via a least-squares problem. This formulation yields improved metric depth accuracy, strong robustness under extreme anchor sparsity, and an interpretable decomposition of spatial scale variations. Extensive experiments across multiple datasets and representative MDE models demonstrate the effectiveness and general applicability of our approach.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。