用游戏分支剧情测试模型对社交隐情的推理能力,发现多数模型只看表面。
SocialReasonBench: A Video-QA Benchmark for Social Reasoning with Counterfactual Narrative Videos

- 基于游戏《底特律:变人》的多分支剧情生成带反事实选项的视频问答题
- 模型在基本社交理解上尚可,但在因果与假设推理上准确率不足50%
- 专设陷阱项揭露模型依赖视觉捷径而非深层社会状态推理
大型多模态模型虽提升了视频理解能力,但在人类中心的社会情境推理上仍显不足。现有基准多依赖单一叙事轨迹,难以区分模型是真正理解社会动态还是仅捕捉常见模式。我们提出SocialReasonBench,一个基于互动叙事的视频多选问答基准,源自《底特律:变人》的游戏玩法视频。该基准利用玩家选择导致的分支剧情,通过游戏脚本、流程图和已录分支验证不同社会结果。我们构建了多智能体标注流程,定位社会意义片段,以游戏状态信号锚定答案标签,并生成基于理论的带诊断性干扰项的问题。基准涵盖七类推理维度,包括意图识别、情感共情、道德困境、反事实推理与因果前因。对主流LMMs的实验显示,模型在基础社交理解上表现尚可,但反事实与因果推理准确率低于50%。消融与诊断分析表明,模型常依赖不完整模态线索,陷入视觉捷径等推理陷阱,暴露了可观测事件识别与深层社会状态推理之间的差距。
原文摘要 · Abstract (English)
Recent advances in Large Multimodal Models (LMMs) have greatly improved video understanding, yet their ability to reason about human-centered social situations remains limited. Existing benchmarks typically rely on videos with a single observed trajectory, making it difficult to determine whether models truly understand social dynamics or merely exploit recurring narrative patterns. We introduce SocialReasonBench, a video multiple-choice QA benchmark for evaluating socially grounded reasoning in scenarios derived from interactive narratives. Built from gameplay videos of Detroit: Become Human, the benchmark leverages branching storylines where player decisions lead to alternative social outcomes that can be checked against the game's own script, flowchart, and recorded branches. We develop a multi-agent curation pipeline that localizes socially meaningful clips, grounds answer labels in game-state signals, and generates theory-guided questions with diagnostic distractors. SocialReasonBench covers seven reasoning dimensions, including intent recognition, emotional empathy, moral dilemma, counterfactual reasoning, and causal antecedent. Experiments on contemporary LMMs show that models perform reasonably well on basic social understanding but struggle with counterfactual and causal reasoning. Further ablation and diagnostic error analyses reveal that models often depend on incomplete modality cues and fall into reasoning traps such as visual shortcuts, highlighting a gap between observable event recognition and deeper reasoning over latent social states.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。