arXiv:2512.21507cs.CV2025-12被引 5

首个评估视频生成模型社会推理能力的基准,揭示现有模型在深层社交理解上的不足。

SVBench: Evaluation of Video Generation Models on Social Reasoning

  • 基于心理学经典范式构建无训练的评估流水线
  • 七款顶尖视频模型在社会推理上普遍表现不佳
  • 适合研究视频生成与认知建模的学者参考

近期文本到视频生成模型在视觉真实感、运动保真度和文本-视频对齐方面取得显著进展,但仍难以生成具有社会一致性行为的内容。与人类能从简短视觉线索中推断意图、信念、情绪和社会规范不同,当前模型常生成字面意义的场景,未能捕捉潜在因果与心理动态。为此,我们提出首个面向视频生成中社会推理的基准(SVBench)。该基准基于发展与社会心理学,涵盖30个经典社会认知范式,覆盖7个核心维度:心智状态推断、目标导向行为、共同注意、社会协调、亲社会行为、社会规范与多智能体策略。通过全训练无关的代理驱动流程,我们提炼各范式的推理结构,合成多样化的视频可用场景,利用基于提示的批判实现概念中立性与难度控制,并采用高容量视觉语言模型(VLM)裁判,在五个可解释的社会推理维度上评估生成视频。基于此框架,我们首次对七款前沿视频生成系统进行了大规模评估。结果表明,表面合理性与深层社会推理之间存在明显差距,提示当前模型在生成社会根基行为方面仍存局限。项目主页:https://github.com/Gloria2tt/SVBench-Evaluation

原文摘要 · Abstract (English)

Recent text-to-video generation models have made remarkable progress in visual realism, motion fidelity, and text-video alignment, yet they still struggle to produce socially coherent behavior. Unlike humans, who readily infer intentions, beliefs, emotions, and social norms from brief visual cues, current models often generate literal scenes without capturing the underlying causal and psychological dynamics. To systematically assess this limitation, we introduce the first benchmark for social reasoning in video generation. Grounded in developmental and social psychology, the benchmark covers thirty classic social cognition paradigms spanning seven core dimensions: mental-state inference, goal-directed action, joint attention, social coordination, prosocial behavior, social norms, and multi-agent strategy. To operationalize these paradigms, we build a fully training-free agent-based pipeline that distills the reasoning structure of each paradigm, synthesizes diverse video-ready scenarios, enforces conceptual neutrality and difficulty control through cue-based critique, and evaluates generated videos with a high-capacity VLM judge along five interpretable dimensions of social reasoning. Using this framework, we conduct the first large-scale evaluation of seven state-of-the-art video generation systems. Results show a clear gap between surface-level plausibility and deeper social reasoning, suggesting that current models remain limited in their ability to generate socially grounded behavior. https://github.com/Gloria2tt/SVBench-Evaluation

视频生成社会推理评估基准认知建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。