用三路摄像头融合头手身姿态,实时判断自动驾驶司机接管意愿。
Driver-Net: Multi-Camera Fusion for Assessing Driver Take-Over Readiness in Automated Vehicles
- 三路摄像头同步捕捉头、手、身体姿态,多视角融合输入。
- 在利兹大学模拟器数据集上达到95.8%的接管就绪分类准确率。
- 适合关注自动驾驶安全交接的工程师与标准制定者。
确保自动驾驶车辆控制权安全转移,需精准及时评估驾驶员准备状态。本文提出Driver-Net,一种新型深度学习框架,通过多摄像头输入融合来估计驾驶员接管就绪程度。与传统基于头部姿态或眼球注视的驾驶监控系统不同,Driver-Net利用三路摄像头同步获取驾驶员头部、双手及身体姿态的视觉线索。模型采用双路径架构,包含上下文模块(Context Block)和特征模块(Feature Block),并结合跨模态融合策略以提升预测精度。在利兹大学驾驶模拟器收集的多样化数据集上评估,该方法在驾驶员就绪状态分类任务中最高达95.8%的准确率。这一性能显著优于现有方法,凸显了多模态与多视角融合的重要性。作为实时、非侵入式解决方案,Driver-Net对构建更安全可靠的自动驾驶系统具有重要意义,并契合新的监管要求与未来安全标准。
原文摘要 · Abstract (English)
Ensuring safe transition of control in automated vehicles requires an accurate and timely assessment of driver readiness. This paper introduces Driver-Net, a novel deep learning framework that fuses multi-camera inputs to estimate driver take-over readiness. Unlike conventional vision-based driver monitoring systems that focus on head pose or eye gaze, Driver-Net captures synchronised visual cues from the driver's head, hands, and body posture through a triple-camera setup. The model integrates spatio-temporal data using a dual-path architecture, comprising a Context Block and a Feature Block, followed by a cross-modal fusion strategy to enhance prediction accuracy. Evaluated on a diverse dataset collected from the University of Leeds Driving Simulator, the proposed method achieves an accuracy of up to 95.8% in driver readiness classification. This performance significantly enhances existing approaches and highlights the importance of multimodal and multi-view fusion. As a real-time, non-intrusive solution, Driver-Net contributes meaningfully to the development of safer and more reliable automated vehicles and aligns with new regulatory mandates and upcoming safety standards.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。