arXiv:2605.15764cs.CVcs.AI2026-05被引 2

构建多人群体非语言互动的社交推理数据集,提升模型对互动对象的识别能力。

GRASP: Learning to Ground Social Reasoning in Multi-Person Non-Verbal Interactions

论文配图:GRASP: Learning to Ground Social Reasoning in Multi-Person Non-Verbal Interactions
图 1 · 摘自论文原文
  • 基于凝视与指代手势轨迹构建社交事件,生成细粒度问答对
  • 在46,000段视频上构建29万组问答,覆盖16类社交推理任务
  • 提出社交对齐奖励机制,显著提升模型对互动主体的推理准确率

理解社交互动需要分析细微的非语言线索,但现有多模态大模型常无法识别多人视频中谁与谁互动。我们提出GRASP,一个大规模社交推理数据集,将高层级社交问题问答与细粒度的凝视和指代手势事件关联。GRASP包含46,000段视频(总计749小时)上的29万组问题-答案对,按16类分类体系组织,涵盖凝视、手势及联合凝视-手势推理。同时提供GRASP-Bench用于评估。与以往仅关注孤立线索或高层问答的数据集不同,GRASP从身份一致的凝视轨迹、指代手势及其组合中构建问题。此外,我们提出社交对齐奖励(SGR),利用这些社交事件作为学习信号,引导模型关注每项互动中的参与者。实验表明,SGR在GRASP-Bench上提升性能,同时保持在相关社交视频问答基准上的零样本表现。

原文摘要 · Abstract (English)

Understanding social interactions requires reasoning over subtle non-verbal cues, yet current multimodal large language models (MLLMs) often fail to identify who interacts with whom in multi-person videos. We introduce GRASP, a large-scale social reasoning dataset that connects high-level social QA with fine-grained gaze and deictic gesture events. GRASP contains 290K question--answer pairs over 46K videos totaling 749 hours, organized by a 16-category taxonomy spanning gaze, gesture, and joint gaze--gesture reasoning, together with GRASP-Bench for evaluation. Unlike prior resources that focus on either isolated cues or high-level social QA, GRASP builds questions from identity-consistent gaze trajectories, deictic gestures, and their joint compositions into social events. Moreover, we propose Social Grounding Reward (SGR), a learning signal that uses these social events to encourage models to reason about the participants involved in each interaction. Experiments show that SGR improves performance on GRASP-Bench while maintaining zero-shot performance on related social video QA benchmarks.

社交推理多模态视觉问答行为识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。