用视线和标记集提升视觉大模型预测用户动作意图
Leveraging Gaze and Set-of-Mark in VLLMs for Human-Object Interaction Anticipation from Egocentric Videos
- 引入标记集提示增强视觉定位能力
- 利用最近注视轨迹理解用户意图,准确率超现有方法
- 适合研究眼动追踪与人机交互的学者
在第一人称视觉中预测人-物交互对智能辅助系统至关重要,有助于引导用户完成日常任务并理解其短期与长期目标。本文针对第一人称视频中的人-物交互预测问题,基于视觉大语言模型(VLLMs)提出新方法。通过引入标记集提示(Set-of-Mark prompting)提升视觉定位能力,并利用用户最近注视点形成的轨迹来理解意图。为捕捉交互前的时序动态,设计了逆指数采样策略处理输入视频帧。在第一人称数据集HD-EPIC上的实验表明,该方法超越现有最先进方法,且具有模型无关性。
原文摘要 · Abstract (English)
The ability to anticipate human-object interactions is highly desirable in an intelligent assistive system in order to guide users during daily life activities and understand their short and long-term goals. Creating systems with such capabilities requires to approach several complex challenges. This work addresses the problem of human-object interaction anticipation in Egocentric Vision using Vision Large Language Models (VLLMs). We tackle key limitations in existing approaches by improving visual grounding capabilities through Set-of-Mark prompting and understanding user intent via the trajectory formed by the user's most recent gaze fixations. To effectively capture the temporal dynamics immediately preceding the interaction, we further introduce a novel inverse exponential sampling strategy for input video frames. Experiments conducted on the egocentric dataset HD-EPIC demonstrate that our method surpasses state-of-the-art approaches for the considered task, showing its model-agnostic nature.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。