arXiv:2601.17895cs.CVcs.RO2026-01被引 25

用视觉上下文修复深度图,让相机在复杂表面也能精准感知空间。

Masked Depth Modeling for Spatial Perception

  • 通过掩码深度建模,利用图像上下文推测缺失或错误的深度信息。
  • 在真实与模拟数据上训练,深度精度超越顶级RGB-D相机。
  • 适合自动驾驶、机器人抓取等需要精确3D感知的场景。

空间视觉感知是自动驾驶和机器人操作等物理世界应用的基础,依赖于对3D环境的交互能力。使用RGB-D相机获取像素级度量深度是最可行的方法,但常受硬件限制和挑战性成像条件影响,尤其是在镜面或无纹理表面。本文认为,深度传感器的误差可视为“被掩码”的信号,反映了潜在的几何模糊性。基于此,我们提出LingBot-Depth深度补全模型,利用视觉上下文通过掩码深度建模优化深度图,并引入自动化数据清洗流程以支持规模化训练。实验表明,该模型在深度精度和像素覆盖率方面均优于顶级RGB-D相机。在多个下游任务中的表现也证明其在RGB与深度模态间提供了对齐的潜在表征。我们公开代码、模型检查点及300万组RGB-深度配对数据(含200万真实数据和100万仿真数据),供空间感知研究社区使用。

原文摘要 · Abstract (English)

Spatial visual perception is a fundamental requirement in physical-world applications like autonomous driving and robotic manipulation, driven by the need to interact with 3D environments. Capturing pixel-aligned metric depth using RGB-D cameras would be the most viable way, yet it usually faces obstacles posed by hardware limitations and challenging imaging conditions, especially in the presence of specular or texture-less surfaces. In this work, we argue that the inaccuracies from depth sensors can be viewed as "masked" signals that inherently reflect underlying geometric ambiguities. Building on this motivation, we present LingBot-Depth, a depth completion model which leverages visual context to refine depth maps through masked depth modeling and incorporates an automated data curation pipeline for scalable training. It is encouraging to see that our model outperforms top-tier RGB-D cameras in terms of both depth precision and pixel coverage. Experimental results on a range of downstream tasks further suggest that LingBot-Depth offers an aligned latent representation across RGB and depth modalities. We release the code, checkpoint, and 3M RGB-depth pairs (including 2M real data and 1M simulated data) to the community of spatial perception.

深度补全空间感知多模态融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。