arXiv:2505.04147cs.CVcs.AI2025-05被引 5

构建复杂社交场景的视频数据集,评估模型读心能力。

R^3-VQA: "Read the Room" by Video Social Reasoning

  • 构建包含信念、意图、欲望、情绪的细粒度社交事件标注数据集
  • 现有大模型在复杂社交推理中仍远低于人类一致性水平
  • 提出基于心理理论的提示方法可提升模型社交推理表现

人类日常生活中具备‘读取情境’的社会推理能力,能从细微社会线索推断他人心理状态。以往社交推理任务与数据集存在场景简单、互动基础、心理变量不全、推理单步等局限,难以模拟真实社交挑战。本文提出高质量、全面的视频数据集 R^3-VQA,包含复杂社交场景下社交事件与心理状态(信念、意图、欲望、情绪)的精确细粒度标注,以及对应的社交因果链。任务涵盖社交事件理解、心理状态估计与社交因果推理三方面。作为基准,我们系统评估了当前主流大视觉语言模型(LVLMs)的社会推理能力与一致性。实验表明:(i) LVLMs 在复杂社交场景中的社会推理仍远未达到人类一致水平;(ii) 使用心理理论(ToM)提示可显著提升模型表现。部分数据与代码已放于附录,全文数据与代码将在录用后公开。

原文摘要 · Abstract (English)

"Read the room" is a significant social reasoning capability in human daily life. Humans can infer others' mental states from subtle social cues. Previous social reasoning tasks and datasets lack complexity (e.g., simple scenes, basic interactions, incomplete mental state variables, single-step reasoning, etc.) and fall far short of the challenges present in real-life social interactions. In this paper, we contribute a valuable, high-quality, and comprehensive video dataset named R^3-VQA with precise and fine-grained annotations of social events and mental states (i.e., belief, intent, desire, and emotion) as well as corresponding social causal chains in complex social scenarios. Moreover, we include human-annotated and model-generated QAs. Our task R^3-VQA includes three aspects: Social Event Understanding, Mental State Estimation, and Social Causal Reasoning. As a benchmark, we comprehensively evaluate the social reasoning capabilities and consistencies of current state-of-the-art large vision-language models (LVLMs). Comprehensive experiments show that (i) LVLMs are still far from human-level consistent social reasoning in complex social scenarios; (ii) Theory of Mind (ToM) prompting can help LVLMs perform better on social reasoning tasks. We provide some of our dataset and codes in supplementary material and will release our full dataset and codes upon acceptance.

社交推理视觉问答多模态心理理论

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。