arXiv:2508.01853cs.CVcs.HC2025-08被引 3

用脑电与眼动数据区分真实场景中目标与非目标注视,准确率达83.6%。

Distinguishing Target and Non-Target Fixations with EEG and Eye Tracking in Realistic Visual Scenes

  • 结合眼动与脑电特征,实现自由视觉搜索中的注视分类
  • 跨用户评估准确率达83.6%,显著优于以往56.9%的基准方法
  • 适用于桌面图标搜索和杂乱车间工具查找等真实场景

在视觉搜索中区分目标与非目标注视是理解用户意图、构建有效辅助系统的基础。以往研究虽证明可基于眼动与脑电(EEG)数据分类目标/非目标注视,但其依赖显式指令、抽象刺激且忽略场景上下文,难以推广至真实场景。本文首次在真实视觉场景中开展自由搜索下的目标/非目标注视分类研究。通过36名参与者、140种真实场景的实验,覆盖桌面图标搜索与杂乱车间工具查找两大应用情境。所提方法融合眼动与EEG特征,超越仅使用注视时长与扫视相关电位的现有最优方法。跨用户评估显示准确率达83.6%,显著高于此前基于扫视相关电位方法的56.9%。实证验证了该方法在不同场景间的泛化能力。

原文摘要 · Abstract (English)

Distinguishing target from non-target fixations during visual search is a fundamental building block to understand users' intended actions and to build effective assistance systems. While prior research indicated the feasibility of classifying target vs. non-target fixations based on eye tracking and electroencephalography (EEG) data, these studies were conducted with explicitly instructed search trajectories, abstract visual stimuli, and disregarded any scene context. This is in stark contrast with the fact that human visual search is largely driven by scene characteristics and raises questions regarding generalizability to more realistic scenarios. To close this gap, we, for the first time, investigate the classification of target vs. non-target fixations during free visual search in realistic scenes. In particular, we conducted a 36-participants user study using a large variety of 140 realistic visual search scenes in two highly relevant application scenarios: searching for icons on desktop backgrounds and finding tools in a cluttered workshop. Our approach based on gaze and EEG features outperforms the previous state-of-the-art approach based on a combination of fixation duration and saccade-related potentials. We perform extensive evaluations to assess the generalizability of our approach across scene types. Our approach significantly advances the ability to distinguish between target and non-target fixations in realistic scenarios, achieving 83.6% accuracy in cross-user evaluations. This substantially outperforms previous methods based on saccade-related potentials, which reached only 56.9% accuracy.

眼动追踪脑电分析视觉搜索人机交互

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。