arXiv:2511.01237cs.CVcs.AI2025-11被引 2

用视线引导视觉注意力,提升第一视角视频的物体检测精度。

Eyes on Target: Gaze-Aware Object Detection in Egocentric Video

  • 将视线特征注入ViT的注意力机制,优先关注人注视区域。
  • 在自建仿真数据集和Ego4D等公开数据集上均显著提升检测准确率。
  • 适合研究人类视觉注意力与智能感知融合的科研人员。

人类视线为理解复杂视觉环境中视觉注意提供了丰富的监督信号。本文提出Eyes on Target,一种面向第一视角视频的深度感知与视线引导的物体检测框架。该方法将视线导出的特征注入视觉变压器(ViT)的注意力机制,有效引导空间特征选择偏向人类关注区域。与传统检测器均匀处理所有区域不同,本方法强调观察者优先区域,从而提升物体检测性能。我们在一个第一视角模拟数据集上验证了该方法,该数据集中人类视觉注意力对任务评估至关重要,展示了其在模拟场景中评估人类表现的潜力。通过大量实验与消融分析,证明了视线融合模型在自建仿真数据集及公共基准(包括Ego4D Ego-Motion和Ego-CH-Gaze)上均持续优于无视线基线。此外,我们引入一种视线感知的注意力头重要性度量,揭示了视线线索如何调节Transformer注意力动态。

原文摘要 · Abstract (English)

Human gaze offers rich supervisory signals for understanding visual attention in complex visual environments. In this paper, we propose Eyes on Target, a novel depth-aware and gaze-guided object detection framework designed for egocentric videos. Our approach injects gaze-derived features into the attention mechanism of a Vision Transformer (ViT), effectively biasing spatial feature selection toward human-attended regions. Unlike traditional object detectors that treat all regions equally, our method emphasises viewer-prioritised areas to enhance object detection. We validate our method on an egocentric simulator dataset where human visual attention is critical for task assessment, illustrating its potential in evaluating human performance in simulation scenarios. We evaluate the effectiveness of our gaze-integrated model through extensive experiments and ablation studies, demonstrating consistent gains in detection accuracy over gaze-agnostic baselines on both the custom simulator dataset and public benchmarks, including Ego4D Ego-Motion and Ego-CH-Gaze datasets. To interpret model behaviour, we also introduce a gaze-aware attention head importance metric, revealing how gaze cues modulate transformer attention dynamics.

第一视角视线引导物体检测视觉注意力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。