arXiv:2512.01707cs.CVcs.AI2025-12中稿 · CVPR被引 8

首个评估模型在流媒体视频中利用眼动追踪进行前瞻理解的基准测试

StreamGaze: Gaze-Guided Temporal Reasoning and Proactive Understanding in Streaming Videos

  • 构建眼动引导的时序与前瞻任务,检验模型对注意力转移的感知能力
  • 多模态大模型在眼动辅助任务上与人类表现存在显著差距
  • 适合研究眼动交互、增强现实和实时视频理解的学者使用

流媒体视频理解要求模型不仅能处理连续输入的帧,还需预判用户意图以支持增强现实眼镜等真实应用场景。现有流媒体评测侧重时序推理,但未衡量多模态大语言模型(MLLMs)在流式场景下对人类眼动信号的解读与利用能力。为此,我们提出首个相关基准StreamGaze,全面评估MLLMs在流媒体视频中利用眼动信息进行时序与前瞻推理的能力。StreamGaze引入眼动引导的过去、现在与前瞻任务,检验模型能否基于实时眼动信号追踪注意力变化,并仅依据历史与当前帧推断用户意图。为构建该基准,我们开发了眼动-视频问答生成流水线,通过注视点提取、区域视觉提示与扫描路径构建,将第一人称视频与原始眼动轨迹对齐,生成时空精准的问答对,反映人类感知动态。在所有任务中,最先进的MLLMs与人类表现之间均存在显著性能差距,揭示了当前在眼动驱动的时序推理、意图建模与前瞻性预测方面的关键局限。我们进一步分析了眼动提示策略、推理行为及任务特定失败模式,为未来研究提供洞见。所有数据与代码均已公开,推动眼动引导的流媒体视频理解研究。

原文摘要 · Abstract (English)

Streaming video understanding requires models not only to process temporally incoming frames, but also to anticipate user intention for realistic applications such as Augmented Reality (AR) glasses. While prior streaming benchmarks evaluate temporal reasoning, none measure whether Multimodal Large Language Models (MLLMs) can interpret or leverage human gaze signals within a streaming setting. To fill this gap, we introduce StreamGaze, the first benchmark designed to evaluate how effectively MLLMs utilize gaze for temporal and proactive reasoning in streaming videos. StreamGaze introduces gaze-guided past, present, and proactive tasks that comprehensively assess streaming video understanding. These tasks evaluate whether models can use real-time gaze signals to follow shifting attention and infer user intentions based only on past and currently observed frames. To build StreamGaze, we develop a gaze-video Question Answering (QA) generation pipeline that aligns egocentric videos with raw gaze trajectories through fixation extraction, region-specific visual prompting, and scanpath construction. This pipeline produces spatio-temporally grounded QA pairs that reflect human perceptual dynamics. Across all StreamGaze tasks, we observe substantial performance gaps between state-of-the-art MLLMs and human performance, highlighting key limitations in gaze-based temporal reasoning, intention modeling, and proactive prediction. We further provide detailed analyses of gaze prompting strategies, reasoning behaviors, and task-specific failure modes, offering insights into current limitations and directions for future research. All data and code are publicly available to support continued research in gaze-guided streaming video understanding.

眼动追踪流媒体理解多模态模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。