arXiv:2604.03667cs.CV2026-04中稿 · International Conf…被引 2

用视线和标记集提升视觉大模型预测用户动作意图

Leveraging Gaze and Set-of-Mark in VLLMs for Human-Object Interaction Anticipation from Egocentric Videos

  • 引入标记集提示增强视觉定位能力
  • 利用最近注视轨迹理解用户意图,准确率超现有方法
  • 适合研究眼动追踪与人机交互的学者

在第一人称视觉中预测人-物交互对智能辅助系统至关重要,有助于引导用户完成日常任务并理解其短期与长期目标。本文针对第一人称视频中的人-物交互预测问题,基于视觉大语言模型(VLLMs)提出新方法。通过引入标记集提示(Set-of-Mark prompting)提升视觉定位能力,并利用用户最近注视点形成的轨迹来理解意图。为捕捉交互前的时序动态,设计了逆指数采样策略处理输入视频帧。在第一人称数据集HD-EPIC上的实验表明,该方法超越现有最先进方法,且具有模型无关性。

原文摘要 · Abstract (English)

The ability to anticipate human-object interactions is highly desirable in an intelligent assistive system in order to guide users during daily life activities and understand their short and long-term goals. Creating systems with such capabilities requires to approach several complex challenges. This work addresses the problem of human-object interaction anticipation in Egocentric Vision using Vision Large Language Models (VLLMs). We tackle key limitations in existing approaches by improving visual grounding capabilities through Set-of-Mark prompting and understanding user intent via the trajectory formed by the user's most recent gaze fixations. To effectively capture the temporal dynamics immediately preceding the interaction, we further introduce a novel inverse exponential sampling strategy for input video frames. Experiments conducted on the egocentric dataset HD-EPIC demonstrate that our method surpasses state-of-the-art approaches for the considered task, showing its model-agnostic nature.

人机交互视觉大模型眼动追踪

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。