利用声音回声提升单目深度估计的精度与尺度准确性
AVS-Net: Audio-Visual Scale Net for Self-supervised Monocular Metric Depth Estimation
- 结合视听信号,用回声作为尺度监督信号
- 在多个基准上提升现有方法的深度预测精度
- 适合需要高精度尺度的自动驾驶与机器人场景
从单目视频中进行度量深度预测面临跨数据集泛化性差的问题,且需标注深度数据以进行尺度校准训练。基于多视角重建的自监督方法虽可利用大规模自然视频,但无法提供正确尺度,限制了其效果。近期研究发现,物体反射的可听回声可用于实现无视觉信号下的尺度重建,因声波传播速度固定,可解决尺度和外观的模糊性。然而,直接端到端从声像信号预测深度无法受益于无需声音标注的无监督深度预测方法。本文展示回声在两个方面的作用:一是作为监督信号用于从有标注数据学习度量深度;二是作为尺度校准信号用于自监督训练。实验表明,该方法可提升多个先进模型的性能,并能对自监督深度模型实现尺度校准。
原文摘要 · Abstract (English)
Metric depth prediction from monocular videos suffers from bad generalization between datasets and requires supervised depth data for scale-correct training. Self-supervised training using multi-view reconstruction can benefit from large scale natural videos but not provide correct scale, limiting its benefits. Recently, reflecting audible Echoes off objects is investigated for improved depth prediction and was shown to be sufficient to reconstruct objects at scale even without a visual signal. Because Echoes travel at fixed speed, they have the potential to resolve ambiguities in object scale and appearance. However, predicting depth end-to-end from sound and vision cannot benefit from unsupervised depth prediction approaches, which can process large scale data without sound annotation. In this work we show how Echoes can benefit depth prediction in two ways: When learning metric depth learned from supervised data and as supervisory signal for scale-correct self-supervised training. We show how we can improve the predictions of several state-of-the-art approaches and how the method can scale-correct a self-supervised depth approach.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。