通过眼神与视觉场景的互动,识别司机的注意力分散。
EyeCue: Driver Cognitive Distraction Detection via Gaze-Empowered Egocentric Video Understanding

- 融合眼动与第一视角视频,建模注意力动态变化
- 在多场景数据集上达到74.38%准确率,优于11个基线
- 适用于不同道路、时段和天气,适合智能驾驶安全研究
驾驶员认知分心是道路事故的主要原因,且难以检测。与手动或视觉分心不同,认知分心表现为思维偏离驾驶任务,即使司机外表专注且无明显肢体动作。本文提出EyeCue,一个基于眼动的第一视角视频理解框架,用于检测认知分心。核心洞察是:认知分心体现于眼动与视觉上下文的交互。为捕捉这一交互,EyeCue将眼动信息与第一视角视频结合,实现对司机注意力随时间演变的上下文感知建模。此外,针对现有数据集规模小、多样性不足的问题,本文构建了CogDrive数据集,通过在四个现有驾驶数据集基础上添加认知分心标注,实现多场景覆盖。在CogDrive上的大量实验表明,EyeCue取得74.38%最高准确率,比6种模型家族的11个基线高出7%以上。值得注意的是,其在不同道路类型、时段和天气条件下均保持超70%准确率,展现出强泛化能力。结果凸显了建模眼动-上下文交互的重要性,以及跨模态交互在多模态认知分心检测中的有效性。代码与数据集资源已开源。
原文摘要 · Abstract (English)
Driver cognitive distraction is a major cause of road collisions and remains difficult to detect. Unlike manual or visual distraction, cognitive distraction is diverted by thoughts unrelated to driving, even when the driver appears visually attentive and exhibits no explicit physical movements. In this work, we propose EyeCue, a gaze-empowered egocentric video understanding framework, to detect driver cognitive distraction. A key insight is that cognitive distraction manifests in the interaction between eye gaze and visual context. To capture this interaction, EyeCue integrates eye gaze with egocentric video to enable context-aware modeling of the driver's attention over time. Furthermore, to tackle the limited scale and diversity of existing datasets, we introduce CogDrive, a comprehensive multi-scenario dataset that augments four existing driving datasets with cognitive distraction annotations. Through extensive evaluations on CogDrive, we show that EyeCue achieves the highest accuracy of 74.38%, outperforming 11 baselines from 6 model families by over 7%. Notably, EyeCue can achieve an accuracy of over 70% across various driving scenarios (different road types, times of day, and weather conditions) with strong generalizability. These results highlight the importance of modeling gaze-context interactions and the effectiveness of cross-modal interaction modeling for multimodal cognitive distraction detection. Our codes and CogDrive dataset resources are available at https://github.com/langzhang2000/EyeCue.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。