用声音和视觉融合实现低成本动态行人感知
AV-PedAware: Self-Supervised Audio-Visual Fusion for Dynamic Pedestrian Awareness
- 通过自监督学习融合音频与视觉信息,模拟人类感知
- 在极端光照下仍能实现接近激光雷达的3D行人检测效果
- 适合预算有限但需鲁棒感知的机器人应用
本文提出AV-PedAware,一种基于自监督学习的音视频融合系统,用于提升机器人对动态行人的感知能力。传统依赖摄像头与激光雷达的方法成本高,易受光照、遮挡和天气影响。本方案利用低成像成本的音视频融合,首次尝试通过监听脚步声来预测周边行人的运动轨迹。系统基于激光雷达生成的标签进行自监督训练,无需昂贵标注。实验表明,仅使用音频与视觉数据,即可在极端视觉条件下实现可靠的3D行人检测,性能接近激光雷达系统。研究还构建了新的多模态行人检测数据集,并将公开代码与数据供社区使用,推动机器人感知发展。
原文摘要 · Abstract (English)
In this study, we introduce AV-PedAware, a self-supervised audio-visual fusion system designed to improve dynamic pedestrian awareness for robotics applications. Pedestrian awareness is a critical requirement in many robotics applications. However, traditional approaches that rely on cameras and LIDARs to cover multiple views can be expensive and susceptible to issues such as changes in illumination, occlusion, and weather conditions. Our proposed solution replicates human perception for 3D pedestrian detection using low-cost audio and visual fusion. This study represents the first attempt to employ audio-visual fusion to monitor footstep sounds for the purpose of predicting the movements of pedestrians in the vicinity. The system is trained through self-supervised learning based on LIDAR-generated labels, making it a cost-effective alternative to LIDAR-based pedestrian awareness. AV-PedAware achieves comparable results to LIDAR-based systems at a fraction of the cost. By utilizing an attention mechanism, it can handle dynamic lighting and occlusions, overcoming the limitations of traditional LIDAR and camera-based systems. To evaluate our approach's effectiveness, we collected a new multimodal pedestrian detection dataset and conducted experiments that demonstrate the system's ability to provide reliable 3D detection results using only audio and visual data, even in extreme visual conditions. We will make our collected dataset and source code available online for the community to encourage further development in the field of robotics perception systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。