arXiv:2603.23190cs.CV2026-03

用眼球注视信息增强视觉语言模型,提升对第一人称行为的预测能力

Gaze-Regularized VLMs for Ego-Centric Behavior Understanding

  • 在训练中直接融入眼球注视数据,生成聚焦区域的动态查询
  • 相比不使用注视信息的基线模型,语义评分提升近13%
  • 适合需要精准未来行为预测的应用场景

眼球注视(包括注视点和快速扫视)提供了关于人类意图和未来动作的关键线索。本文提出一种注视正则化框架,用于增强视觉语言模型(VLMs)在第一人称行为理解中的表现。与仅依赖视觉数据而忽略注视信息的现有方法不同,本方法在训练阶段直接将注视信息融入VLM架构。通过生成基于注视的查询,模型能动态关注注视突出的区域;同时,注视正则化机制确保模型注意力与人类注意力模式对齐。为探究注视信息如何有效整合到VLM中,我们开展了大量实验,测试了多种融合策略。这些创新使模型能够生成详细的动作描述以预测未来事件。实验结果表明,相较于未利用注视数据的基线模型,本方法在语义评分上提升近13%,验证了其有效性。本工作为在VLM中利用人类注视信息奠定了基础,显著提升了其在需准确、鲁棒未来事件预测应用中的表现。

原文摘要 · Abstract (English)

Eye gaze, encompassing fixations and saccades, provides critical insights into human intentions and future actions. This study introduces a gaze-regularized framework that enhances Vision Language Models (VLMs) for egocentric behavior understanding. Unlike existing methods that rely solely on visual data and overlook gaze information, our approach directly incorporates gaze information into the VLM architecture during training. By generating gaze-based queries, the model dynamically focuses on gaze-highlighted regions, while a gaze-regularization mechanism ensures the alignment of model attention with human attention patterns. To better understand how gaze can be effectively integrated into VLMs, we conducted extensive experiments exploring various strategies for incorporating gaze data. These innovations enable the prediction of future events with detailed action descriptions. Experimental results demonstrate a nearly 13 % improvement in semantic scores compared to baseline models not leveraging gaze data, highlighting the effectiveness of our approach. This work establishes a foundation for leveraging the human gaze in VLMs, significantly boosting their predictive capabilities in applications requiring accurate and robust future event prediction.

视觉语言模型眼球注视行为预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。