构建社交理解评估框架,测试多模态模型对人类互动的洞察力。
Social Caption: Evaluating Social Understanding in Multimodal Models

- 基于互动理论设计三维度评估框架:社会推断、整体分析、定向分析。
- 实验证明模型规模与架构设计显著影响社交理解能力,语音上下文提升表现。
- 提供可扩展的自动化评估路径,适合研究多模态社会认知的学者使用。
社会理解能力对于多模态大语言模型(MLLMs)解读人类社会互动至关重要。我们提出SOCIAL CAPTION框架,基于互动理论,从三个维度评估MLLMs的社会理解能力:社会推断(SI),即准确推断互动内容的能力;整体社会分析(HSA),即生成对互动的全面描述能力;定向社会分析(DSA),即从互动中提取相关信息的能力。我们分析了影响模型表现的因素,如模型规模、架构设计和语音上下文。通过使用MLLM裁判进行实验,展示了实现多模态社会理解自动化评估的可行路径。
原文摘要 · Abstract (English)
Social understanding abilities are crucial for multimodal large language models (MLLMs) to interpret human social interactions. We introduce SOCIAL CAPTION, a framework grounded in interaction theory to evaluate social understanding abilities of MLLMs along three dimensions: Social Inference (SI), the ability to make accurate inferences about interactions; Holistic Social Analysis (HSA), the ability to generate comprehensive descriptions of interactions; Directed Social Analysis (DSA), the ability to generate relevant information from interactions. We analyze factors influencing model performance in social understanding, such as scale, architectural design, and spoken context. Experiments with MLLM judges demonstrate a path towards scaling automated evaluation of multimodal social understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。