让小型无人机在本地实时学习深度感知,提升避障能力。
On-Device Self-Supervised Learning of Low-Latency Monocular Depth from Only Events
- 基于对比度最大化实现设备端自监督学习
- 在小型无人机上实现实时低延迟深度估计
- 适合资源受限的飞行机器人实时感知场景
事件相机以毫瓦级功耗实现低延迟感知,非常适合小型敏捷机器人(如微型飞行无人机)。基于对比度最大化的自监督学习无需高频真值数据,可在机器人运行环境中在线学习,极具潜力。然而,设备端在线学习面临计算效率与实时性双重挑战。本文优化了对比度最大化流程的时间与内存效率,首次实现小型无人机上的实时深度感知学习。实验表明,在线学习相比仅预训练可显著提升深度估计精度和避障成功率。基准测试显示,该方法在自监督方案中达到顶尖性能。本工作挖掘了设备端在线学习的潜力,有望缩小真实环境差距并提升系统表现。
原文摘要 · Abstract (English)
Event cameras provide low-latency perception for only milliwatts of power. This makes them highly suitable for resource-restricted, agile robots such as small flying drones. Self-supervised learning based on contrast maximization holds great potential for event-based robot vision, as it foregoes the need for high-frequency ground truth and allows for online learning in the robot's operational environment. However, online, on-board learning raises the major challenge of achieving sufficient computational efficiency for real-time learning, while maintaining competitive visual perception performance. In this work, we improve the time and memory efficiency of the contrast maximization pipeline, making on-device learning of low-latency monocular depth possible. We demonstrate that online learning on board a small drone yields more accurate depth estimates and more successful obstacle avoidance behavior compared to only pre-training. Benchmarking experiments show that the proposed pipeline is not only efficient, but also achieves state-of-the-art depth estimation performance among self-supervised approaches. Our work taps into the unused potential of online, on-device robot learning, promising smaller reality gaps and better performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。