用距离变换增强低纹理区轮廓信息,提升单目深度估计精度
Improved monocular depth prediction using distance transform over pre-semantic contours with self-supervised neural networks

- 在预语义轮廓上应用距离变换,强化低纹理区域的空间信息
- 在KITTI等5个数据集上超越现有自监督方法,深度预测更准确
- 适合做自动驾驶、机器人视觉中的单目深度建模任务
基于自监督训练的单目深度估计在低纹理区域表现不佳,因光度损失可能导致模糊的深度预测。为此,我们提出一种新方法:在预语义轮廓上应用距离变换,增强空间信息,提升低纹理区域的判别能力。该方法联合估计预语义轮廓、深度和自身运动。利用预语义轮廓生成新输入图像,通过距离变换在均匀区域增加方差,从而构建更有效的损失函数,改善深度与自身运动的训练效果。理论上证明,在此场景下距离变换是最优的方差增强技术。在KITTI、Cityscapes、Waymo、NYUv2和ScanNet上进行大量实验,结果表明模型表现稳健,显著优于现有自监督方法。
原文摘要 · Abstract (English)
Monocular depth estimation (MDE) with self-supervised training approaches struggles in low-texture areas, where photometric losses may lead to ambiguous depth predictions. To address this, we propose a novel technique that enhances spatial information by applying a distance transform over pre-semantic contours, augmenting discriminative power in low texture regions. Our approach jointly estimates pre-semantic contours, depth and ego-motion. The pre-semantic contours are leveraged to produce new input images, with variance augmented by the distance transform in uniform areas. This approach results in more effective loss functions, enhancing the training process for depth and ego-motion. We demonstrate theoretically that the distance transform is the optimal variance-augmenting technique in this context. Through extensive experiments on KITTI, Cityscapes, Waymo, NYUv2 and ScanNet our model demonstrates robust performance, surpassing competing self-supervised methods in MDE.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。