用视线引导提问,提升智能助手理解日常视频意图的能力
In the Eye of MLLM: Benchmarking Egocentric Video Intent Understanding with Gaze-Guided Prompting
- 利用用户视线信息构建问答数据集,引导模型关注关键区域
- 视线引导使模型在长视频中意图识别准确率显著提升
- 适合研究眼动追踪与多模态大模型融合的开发者
多模态大语言模型(MLLM)的兴起显著提升了AI助手跨模态处理复杂信息的能力。近期,第一人称视频通过直接捕捉用户注意力、行为和上下文,为基于MLLM的主动式、个性化用户体验提供了新可能。然而,现有基准忽略了视线作为用户意图指标的关键作用。为此,我们提出EgoGazeVQA——一个基于视线引导的第一人称视频问答基准,利用视线信息增强对长时日常视频的理解。该数据集包含由MLLM生成并经人工标注优化的视线-问答对。实验表明,现有MLLM难以准确解析用户意图;而我们的视线引导提示方法通过整合空间、时间与意图线索,显著提升性能。我们还开展视线相关微调实验,分析视线估计精度对提示效果的影响。结果表明,视线信息对实现更个性化、高效的沉浸式场景下AI助手具有重要价值。
原文摘要 · Abstract (English)
The emergence of advanced multimodal large language models (MLLMs) has significantly enhanced AI assistants' ability to process complex information across modalities. Recently, egocentric videos, by directly capturing user focus, actions, and context in an unified coordinate, offer an exciting opportunity to enable proactive and personalized AI user experiences with MLLMs. However, existing benchmarks overlook the crucial role of gaze as an indicator of user intent. To address this gap, we introduce EgoGazeVQA, an egocentric gaze-guided video question answering benchmark that leverages gaze information to improve the understanding of longer daily-life videos. EgoGazeVQA consists of gaze-based QA pairs generated by MLLMs and refined by human annotators. Our experiments reveal that existing MLLMs struggle to accurately interpret user intentions. In contrast, our gaze-guided intent prompting methods significantly enhance performance by integrating spatial, temporal, and intent-related cues. We further conduct experiments on gaze-related fine-tuning and analyze how gaze estimation accuracy impacts prompting effectiveness. These results underscore the value of gaze for more personalized and effective AI assistants in egocentric settings. Project page: https://taiyi98.github.io/projects/EgoGazeVQA
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。