arXiv:2409.05587cs.CV2024-09被引 6

融合Transformer与Mamba,提升驾驶分心识别精度与鲁棒性

DSDFormer: An Innovative Transformer-Mamba Framework for Robust High-Precision Driver Distraction Identification

  • 用双状态域注意力融合全局与局部特征
  • 在三个数据集上达最优性能,支持实时推理
  • 适合智能交通系统中的高精度驾驶监控场景

驾驶分心仍是全球交通事故的主要原因,威胁道路安全。随着智能交通系统的发展,实现准确、实时的驾驶分心识别变得至关重要。然而,现有方法难以同时捕捉全局上下文与细粒度局部特征,且面临训练数据中噪声标签的问题。为此,我们提出DSDFormer,一种通过双状态域注意力(DSDA)机制融合Transformer与Mamba架构优势的新框架,平衡长程依赖与细节特征提取,提升驾驶行为识别鲁棒性。此外,引入无监督的时序推理置信学习(TRCL),利用视频序列中的时空相关性优化噪声标签。模型在AUC-V1、AUC-V2和100-Driver数据集上达到当前最优表现,并在NVIDIA Jetson AGX Orin平台实现实时处理。大量实验表明,DSDFormer与TRCL显著提升驾驶分心检测的准确率与鲁棒性,为提升道路安全提供可扩展解决方案。

原文摘要 · Abstract (English)

Driver distraction remains a leading cause of traffic accidents, posing a critical threat to road safety globally. As intelligent transportation systems evolve, accurate and real-time identification of driver distraction has become essential. However, existing methods struggle to capture both global contextual and fine-grained local features while contending with noisy labels in training datasets. To address these challenges, we propose DSDFormer, a novel framework that integrates the strengths of Transformer and Mamba architectures through a Dual State Domain Attention (DSDA) mechanism, enabling a balance between long-range dependencies and detailed feature extraction for robust driver behavior recognition. Additionally, we introduce Temporal Reasoning Confident Learning (TRCL), an unsupervised approach that refines noisy labels by leveraging spatiotemporal correlations in video sequences. Our model achieves state-of-the-art performance on the AUC-V1, AUC-V2, and 100-Driver datasets and demonstrates real-time processing efficiency on the NVIDIA Jetson AGX Orin platform. Extensive experimental results confirm that DSDFormer and TRCL significantly improve both the accuracy and robustness of driver distraction detection, offering a scalable solution to enhance road safety.

驾驶分心识别TransformerMamba实时系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。