arXiv:2502.03005cs.CVcs.LG2025-02被引 2

融合视频音频多模态数据,提升驾驶异常检测准确率。

Driver Assistance System Based on Multimodal Data Hazard Detection

  • 用注意力机制融合道路视频、人脸视频和音频数据
  • 在新构建的三模态数据集上显著降低误判率
  • 适合自动驾驶安全系统研发人员参考

自动驾驶技术已取得显著进展,但因驾驶事件呈现长尾分布,异常检测仍是重大挑战。现有方法主要依赖单一模态的道路视频数据,难以捕捉罕见且不可预测的驾驶事故。本文提出一种融合道路视频、驾驶员面部视频和音频数据的多模态驾驶辅助检测系统。模型采用基于注意力的中间融合策略,实现端到端学习,无需单独特征提取。为支持该方法,我们使用驾驶模拟器构建了一个新的三模态数据集。实验结果表明,该方法能有效捕捉跨模态相关性,减少误判,提升驾驶安全性。

原文摘要 · Abstract (English)

Autonomous driving technology has advanced significantly, yet detecting driving anomalies remains a major challenge due to the long-tailed distribution of driving events. Existing methods primarily rely on single-modal road condition video data, which limits their ability to capture rare and unpredictable driving incidents. This paper proposes a multimodal driver assistance detection system that integrates road condition video, driver facial video, and audio data to enhance incident recognition accuracy. Our model employs an attention-based intermediate fusion strategy, enabling end-to-end learning without separate feature extraction. To support this approach, we develop a new three-modality dataset using a driving simulator. Experimental results demonstrate that our method effectively captures cross-modal correlations, reducing misjudgments and improving driving safety.

多模态驾驶安全异常检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。