引入声音感知,让汽车更懂司机与环境,提升安全决策能力。
Looking and Listening Inside and Outside: Multimodal Artificial Intelligence Systems for Driver Safety Assessment and Intelligent Vehicle Decision-Making
- 融合视听信号构建听看内外框架,增强驾驶状态理解
- 语音分析可识别酒驾等异常状态,提升安全预警能力
- 适合自动驾驶、智能座舱等需多模态感知的场景
传统的‘看内看外’(LILO)框架已用于智能车辆中理解外部环境与驾驶员状态,如智能安全气囊部署、自动驾驶接管时间预测和注意力监测。本文提出扩展该框架,引入音频模态作为补充信息源,形成‘听看内外’(L-LIO)新框架,通过多模态传感器融合提升驾驶员状态评估与环境理解能力。研究评估了三个应用场景:利用驾驶员语音进行受控学习以分类潜在失能状态(如醉酒);采集并分析乘客自然语言指令(如“在那栋红房子后转弯”),推动语音指令与规划系统对接;在视觉系统无法明确时,音频可澄清外部人员的引导或手势。数据集包含真实环境中采集的车内及外部音频样本。初步结果表明,音频在复杂或情境丰富的场景中提供关键安全洞察,尤其当仅依赖视觉信号时效果不足。挑战包括环境噪声干扰、隐私问题以及跨个体鲁棒性,亟需进一步研究动态真实场景下的可靠性。L-LIO通过音视频融合增强了对驾驶员与场景的理解,为安全干预提供了新路径。
原文摘要 · Abstract (English)
The looking-in-looking-out (LILO) framework has enabled intelligent vehicle applications that understand both the outside scene and the driver state to improve safety outcomes, with examples in smart airbag deployment, takeover time prediction in autonomous control transitions, and driver attention monitoring. In this research, we propose an augmentation to this framework, making a case for the audio modality as an additional source of information to understand the driver, and in the evolving autonomy landscape, also the passengers and those outside the vehicle. We expand LILO by incorporating audio signals, forming the looking-and-listening inside-and-outside (L-LIO) framework to enhance driver state assessment and environment understanding through multimodal sensor fusion. We evaluate three example cases where audio enhances vehicle safety: supervised learning on driver speech audio to classify potential impairment states (e.g., intoxication), collection and analysis of passenger natural language instructions (e.g., "turn after that red building") to motivate how spoken language can interface with planning systems through audio-aligned instruction data, and limitations of vision-only systems where audio may disambiguate the guidance and gestures of external agents. Datasets include custom-collected in-vehicle and external audio samples in real-world environments. Pilot findings show that audio yields safety-relevant insights, particularly in nuanced or context-rich scenarios where sound is critical to safe decision-making or visual signals alone are insufficient. Challenges include ambient noise interference, privacy considerations, and robustness across human subjects, motivating further work on reliability in dynamic real-world contexts. L-LIO augments driver and scene understanding through multimodal fusion of audio and visual sensing, offering new paths for safety intervention.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。