arXiv:2606.26922cs.ROcs.AI2026-06

提出轻量级多模态驾驶状态监测框架,实现低延迟安全决策。

Risk-Aware Selective Multimodal Driver Monitoring with Driver-State World Modeling

论文配图:Risk-Aware Selective Multimodal Driver Monitoring with Driver-State World Modeling
图 1 · 摘自论文原文
  • 基于视觉与生理信号的轻量学生模型,结合自适应门控机制选择性推理
  • 将不安全误判率从17.37%降至约5%,推理延迟仅3.08ms
  • 引入驾驶状态世界建模预测未来风险,适合边缘部署的安全系统

自动驾驶车辆中持续的驾驶员监控需在低延迟下避免不确定状态下的危险决策。大型视觉语言模型虽具广泛多模态先验,但其高延迟和可靠性不足,不适合作为始终在线的座舱监控。本文提出一种成本感知的选择性推理框架,核心为轻量级RGB-生理学学生模型,融合车内视觉与窗级心率/皮肤电活动信号,并通过学习到的门控机制决定是否采纳快速预测或放弃以进行安全干预。额外控制显示,学习得分包含样本级信息,超出场景先验,但精确生理同步仍受限。为进一步引入预测证据,研究了紧凑的驾驶状态世界建模模块,可滚动推演潜在驾驶状态特征,估计未来快速模型误差与反事实系统级行动成本。在情景诱导的驾驶需求识别任务上,该模型优于仅视觉或仅生理基线,达到0.7440宏平均F1和0.9099平衡准确率,参数量11.39M,推理延迟3.08ms。成本感知选择性推理将不安全误判率从始终快速推理下的17.37%降低至约5%(跨种子),同时保持部署级延迟。尽管驾驶状态世界建模提供有价值预测信号,最坏组评估揭示持续存在的工作点校准漂移问题。最终,可靠的边缘驾驶监控不仅需提升感知主干,还需风险感知的选择性控制与群体鲁棒校准。

原文摘要 · Abstract (English)

Continuous driver monitoring in automated vehicles requires low-latency inference while avoiding unsafe decisions under uncertain driver states. Large vision-language models provide broad multimodal priors, but their latency and limited reliability in this setting make them unsuitable as always-on in-cabin monitors. We propose a cost-aware selective inference framework for deployable multimodal driver monitoring. The core system is a lightweight RGB-physiological student that combines in-cabin visual observations with window-level HR/EDA signals, and a learned gate that decides when to accept the fast prediction or abstain for safety intervention. Additional controls show that the learned scores contain sample-level information beyond scenario priors, while exact physiological synchronization remains a limitation. To incorporate predictive evidence, we further study a compact driver-state world modeling module that rolls out latent driver-state features and estimates future fast-model errors and counterfactual system-level action costs. On scenario-induced driver-demand recognition, the RGB-physiological student improves over RGB-only and physiology-only baselines, reaching 0.7440 Macro-F1 and 0.9099 balanced accuracy with 11.39M parameters and 3.08ms inference latency. Cost-aware selective inference reduces unsafe false negatives from 17.37% under always-fast inference to approximately 5% across seeds, while maintaining deployment-level latency. While driver-state world modeling offers valuable predictive signals, worst-group evaluations highlight persistent operating-point calibration drift. Ultimately, reliable edge driver monitoring requires advancing not only perception backbones, but also risk-aware selective control and group-robust calibration.

驾驶监控多模态感知边缘计算风险控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。