用事件相机捕捉眼动数据,精准识别用户认知负荷等级。
EveLoad: Cognitive Workload Recognition from Event-Based Eye Movements

- 基于事件相机构建时空融合的眼动特征编码框架。
- 在六级负荷下实现96.36%的个体准确率。
- 首个面向任务驱动场景的事件眼动负荷数据集,适合康复与人机交互研究。
认知负荷监测对自适应康复和辅助界面至关重要,需根据用户认知状态动态调整任务难度、节奏与反馈,避免过载或挑战不足。新兴扩展现实与机器人辅助康复环境提供可控训练任务,但需无侵入式传感方法以捕捉交互中的快速眼动。现有眼动负荷识别多依赖帧基眼动仪,存在时间分辨率低、快速眼动下鲁棒性差的问题。相比之下,事件相机具备微秒级时间分辨率、高动态范围和低延迟,适合捕捉精细眼动动态。以往研究多采用自由注视等范式,注视位置随任务变化,模型可能学习到注视分布与负荷的关联,而非负荷相关眼动特征本身。本文提出EveLoad,据我们所知是首个带有分级认知负荷标注的事件基眼动数据集,基于20名健康参与者在空间受限、任务驱动的N-back引导固定注视范式下采集。基于该数据集,建立六级负荷识别基准,并提出编码时空事件表示的学习框架。实验表明,该方法在混合随机划分评估下达到96.36%和96.13%的平均个体准确率,表明事件基眼动可为未来负荷感知康复系统提供有效传感路径。
原文摘要 · Abstract (English)
Cognitive workload monitoring is important for adaptive rehabilitation and assistive interfaces, where task difficulty, pacing, and feedback should be adjusted according to the user's cognitive state to avoid overload and under-challenge. Emerging extended reality and robot-assisted rehabilitation environments provide controllable training tasks, but they require unobtrusive sensing methods that can capture rapid ocular dynamics during interaction. Existing eye-movement-based cognitive workload recognition methods mainly rely on frame-based eye trackers, which often suffer from limited temporal resolution and degraded robustness under rapid eye movements. In contrast, event cameras provide microsecond-level temporal resolution, high dynamic range and low latency, making them suitable for capturing fine-grained ocular dynamics. Many previous studies rely on free-viewing or similar paradigms, where gaze locations can vary across tasks. As a result, models may learn associations between gaze-location distributions and cognitive workload, rather than workload-related eye movement characteristics themselves. In this work, we introduce EveLoad, which, to the best of our knowledge, is the first event-based eye-movement dataset with graded cognitive workload annotations, collected from 20 healthy participants under spatially constrained and task-driven conditions using a controlled N-back-guided fixation paradigm. Based on this dataset, we establish a benchmark for cognitive workload recognition with six workload levels and propose a learning framework that encodes spatiotemporal event representations. Experimental results show that our approach achieves an average subject-specific accuracy of 96.36% and 96.13% under mixed random split evaluation. These results suggest that event-based eye movements may provide a useful sensing pathway for future workload-aware rehabilitation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。