动态场景下自监督深度估计新方法,有效抑制移动物体干扰
D$^3$epth: Self-Supervised Depth Estimation with Dynamic Mask in Dynamic Scenes
- 用动态掩码识别并抑制动态物体影响
- 多帧融合时通过成本体积自动掩码提升精度
- 引入频谱熵评估不确定性,适合真实动态环境
深度估计在机器人领域至关重要。近期自监督深度估计方法展现出巨大潜力,可高效利用大量无标签真实数据。然而,现有方法大多基于静态场景假设,难以适应动态环境。为此,我们提出 D³epth,一种面向动态场景的自监督深度估计新方法。该方法从两个关键角度应对动态物体挑战:首先,在自监督框架内设计重投影约束,识别可能含动态物体的区域,构建动态掩码以在损失层面减轻其影响;其次,针对多帧深度估计,提出成本体积自动掩码策略,利用相邻帧识别动态区域并生成掩码,为后续过程提供指导;此外,提出频谱熵不确定性模块,引入频谱熵引导深度融合中的不确定性估计,有效缓解动态环境下成本体积计算带来的问题。在 KITTI 与 Cityscapes 数据集上的大量实验表明,所提方法持续优于现有自监督单目深度估计基线。代码已开源。
原文摘要 · Abstract (English)
Depth estimation is a crucial technology in robotics. Recently, self-supervised depth estimation methods have demonstrated great potential as they can efficiently leverage large amounts of unlabelled real-world data. However, most existing methods are designed under the assumption of static scenes, which hinders their adaptability in dynamic environments. To address this issue, we present D$^3$epth, a novel method for self-supervised depth estimation in dynamic scenes. It tackles the challenge of dynamic objects from two key perspectives. First, within the self-supervised framework, we design a reprojection constraint to identify regions likely to contain dynamic objects, allowing the construction of a dynamic mask that mitigates their impact at the loss level. Second, for multi-frame depth estimation, we introduce a cost volume auto-masking strategy that leverages adjacent frames to identify regions associated with dynamic objects and generate corresponding masks. This provides guidance for subsequent processes. Furthermore, we propose a spectral entropy uncertainty module that incorporates spectral entropy to guide uncertainty estimation during depth fusion, effectively addressing issues arising from cost volume computation in dynamic environments. Extensive experiments on KITTI and Cityscapes datasets demonstrate that the proposed method consistently outperforms existing self-supervised monocular depth estimation baselines. Code is available at \url{https://github.com/Csyunling/D3epth}.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。