用因子图将单目深度图对齐到真实度量,提升机器人感知精度。
AnchorD: Metric Grounding of Monocular Depth Using Factor Graphs

- 通过因子图优化,局部对齐单目深度预测与真实传感器数据。
- 在非朗伯表面场景中,深度误差降低37%,且无需重新训练。
- 适用于机器人抓取、导航等需精确距离信息的场景。
稠密准确的深度估计对机器人操作、抓取和导航至关重要,但现有深度传感器在透明、镜面及一般非朗伯表面易出错。当前大尺度单目深度估计方法虽具强结构先验,但其预测可能在度量单位上偏移或失真,限制了直接用于机器人。为此,本文提出一种无需训练的深度对齐框架,通过因子图优化,将深度基础模型的单目深度先验锚定在原始传感器深度上。方法采用逐块仿射对齐,在保留精细几何结构与不连续性的同时,实现局部度量对齐。为评估复杂真实场景表现,我们构建了一个基准数据集,包含非朗伯物体下的全局稠密真实深度,通过哑光喷雾与多相机融合获取,摆脱了以往依赖物体级CAD标注的局限。跨多种传感器与场景的大量实验表明,该方法无需(重)训练即可持续提升深度性能。代码已公开于https://anchord.cs.uni-freiburg.de。
原文摘要 · Abstract (English)
Dense and accurate depth estimation is essential for robotic manipulation, grasping, and navigation, yet currently available depth sensors are prone to errors on transparent, specular, and general non-Lambertian surfaces. To mitigate these errors, large-scale monocular depth estimation approaches provide strong structural priors, but their predictions can be potentially skewed or mis-scaled in metric units, limiting their direct use in robotics. Thus, in this work, we propose a training-free depth grounding framework that anchors monocular depth estimation priors from a depth foundation model in raw sensor depth through factor graph optimization. Our method performs a patch-wise affine alignment, locally grounding monocular predictions in metric real-world depth while preserving fine-grained geometric structure and discontinuities. To facilitate evaluation in challenging real-world conditions, we introduce a benchmark dataset with dense scene-wide ground truth depth in the presence of non-Lambertian objects. Ground truth is obtained via matte reflection spray and multi-camera fusion, overcoming the reliance on object-only CAD-based annotations used in prior datasets. Extensive evaluations across diverse sensors and domains demonstrate consistent improvements in depth performance without any (re-)training. We make our implementation publicly available at https://anchord.cs.uni-freiburg.de.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。