构建视频社交理解基准,评估大模型在社交认知与预测上的短板。
SIV-Bench: A Video Benchmark for Social Interaction Understanding and Reasoning
- 基于社会关系理论设计多模态视频评测集,覆盖14类关系
- 顶尖大模型在社交理解上表现尚可,但推理与预测能力薄弱
- 音频与字幕显著提升复杂推理任务表现,适合社交智能研究者
理解社交互动——包括感知多模态细微线索、推断不可见心理状态与关系、动态预测他人行为——是实现人机交互的基础。尽管多模态大语言模型(MLLMs)快速发展,社交互动的丰富性仍阻碍了全面评估其社交能力的基准建设。基于被广泛接受的社会关系理论,我们提出SIV-Bench,一个系统评估MLLM在社交场景理解(SSU)、社交状态推理(SSR)和社交动态预测(SDP)方面能力的新视频基准。该基准包含2,792个原始采集视频片段和5,455个通过人-大模型协作生成的问答对,涵盖14种典型人际关系,覆盖多样视频时长、类型、呈现风格及语言文化背景。全面实验表明,领先MLLM在SSU上表现较好,但在SSR和SDP上仍较弱,关系推断中的系统性混淆是关键瓶颈。深入分析显示,模型表现不佳源于与人类思维不一致及推理深度不足。此外,音频和字幕对推理密集型的SSR和SDP有显著帮助。SIV-Bench为衡量进展、揭示局限、指导未来研究提供了统一测试平台。数据集与代码已公开:https://kfq20.github.io/sivbench。
原文摘要 · Abstract (English)
Understanding social interaction, which encompasses perceiving numerous and subtle multimodal cues, inferring unobservable mental states and relations, and dynamically predicting others' behavior, is the foundation for achieving human-machine interaction. Despite rapid advances in Multimodal Large Language Models (MLLMs), the rich and multifaceted nature of social interaction has hindered the development of benchmarks that holistically evaluate and guide their social interaction abilities. Based on social relation theory, which has been widely regarded as a foundational framework for understanding social behavior, we provide SIV-Bench, a novel video benchmark for systematically evaluating MLLMs' capabilities across Social Scene Understanding (SSU), Social State Reasoning (SSR), and Social Dynamics Prediction (SDP). SIV-Bench features 2,792 originally collected video clips and 5,455 meticulously generated question-answer pairs derived from a human-LLM collaborative pipeline. It covers 14 typical relationships, diverse video lengths, genres, presentation styles, and linguistic and cultural backgrounds. Our comprehensive experiments show that leading MLLMs perform relatively well on SSU but remain weak on SSR and SDP, with the systematic confusion in relation inference as a key bottleneck. An in-depth analysis of the reasoning process attributes MLLMs' suboptimal performance to misalignment with human thoughts and insufficient reasoning depth. Moreover, we find audio and subtitles aid in reasoning-intensive SSR and SDP. Together, SIV-Bench offers a unified testbed to measure progress, expose limitations, and guide future research toward more socially intelligent MLLMs. We release the dataset and code at our project website: https://kfq20.github.io/sivbench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。