arXiv:2603.12533cs.CV2026-03中稿 · CVPR

让AI看懂人指的方向,提升第一视角视频问答能力

Do You See What I Am Pointing At? Gesture-Based Egocentric Video Question Answering

  • 用3D手部关键点生成手势令牌,嵌入模型输入增强空间时间理解
  • 在6个任务上平均准确率达68.1%,比当前最优模型高6.6%
  • 首次构建真实与合成结合的指指点点视频问答数据集,适合视觉-语言研究者

理解并回答基于用户指向手势的问题,是下一代第一视角人工智能助手的关键。然而,现有多模态大模型因缺乏丰富的手势数据且难以从第一视角视频中推断精细指向意图而表现不佳。为此,我们提出EgoPointVQA数据集与基准,包含4000个合成视频和400个真实世界视频,覆盖多个指示性推理任务。在此基础上,我们进一步提出手部意图令牌(HINT),通过现成重建模型提取3D手部关键点生成令牌,并将其插入模型输入,提供明确的空间与时间上下文以解析指向意图。实验表明,该方法在不同主干网络和模型规模下均优于现有方法;其中,HINT-14B在6个任务上平均准确率达到68.1%,超越当前最优模型InternVL3-14B达6.6%。为推动开放研究,我们将公开代码、模型与数据集。项目页面:https://yuuraa.github.io/papers/choi2026egovqa

原文摘要 · Abstract (English)

Understanding and answering questions based on a user's pointing gesture is essential for next-generation egocentric AI assistants. However, current Multimodal Large Language Models (MLLMs) struggle with such tasks due to the lack of gesture-rich data and their limited ability to infer fine-grained pointing intent from egocentric video. To address this, we introduce EgoPointVQA, a dataset and benchmark for gesture-grounded egocentric question answering, comprising 4000 synthetic and 400 real-world videos across multiple deictic reasoning tasks. Built upon it, we further propose Hand Intent Tokens (HINT), which encodes tokens derived from 3D hand keypoints using an off-the-shelf reconstruction model and interleaves them with the model input to provide explicit spatial and temporal context for interpreting pointing intent. We show that our model outperforms others in different backbones and model sizes. In particular, HINT-14B achieves 68.1% accuracy, on average over 6 tasks, surpassing the state-of-the-art, InternVL3-14B, by 6.6%. To further facilitate the open research, we will release the code, model, and dataset. Project page: https://yuuraa.github.io/papers/choi2026egovqa

第一视角手势理解视觉问答多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。