arXiv:2606.23092cs.CL2026-06

首个评估多模态模型细粒度人际互动推理能力的基准

PIVOTSBench: Evaluating Fine-Grained Interpersonal Relationship Reasoning in Multimodal Large Language Models

论文配图:PIVOTSBench: Evaluating Fine-Grained Interpersonal Relationship Reasoning in Multimodal Large Language Models
图 1 · 摘自论文原文
  • 基于社交智商2.0和YouTube数据构建新基准
  • 模型在双向关系维度预测上表现有限,视觉线索利用不足
  • 适合研究社会智能与多模态推理的学者使用

人类具备理解细粒度人际互动关系的天然能力,这在日常社交中至关重要。尽管此类推理本质上是多模态的,但现有多模态大语言模型(MLLMs)对此研究甚少。为此,我们提出PIVOTS,首个基于Social-IQ 2.0和YouTube数据构建的基准,用于评估MLLMs在心理学研究支持下预测双向人际关系维度的能力。PIVOTS还包含辅助任务,评估模型识别并利用关键视觉线索进行推理的能力。我们评估了专有及开源的MLLMs,并通过详尽消融实验分析了视觉模态与对话中显式社会角色信息的影响。进一步研究联合与成对预测设置对评分双向PIVOTS维度的提升效果。项目主页与资源:https://flynnzhangsx.github.io/PIVOTSBench/

原文摘要 · Abstract (English)

Humans possess an innate ability to understand fine-grained interpersonal relationships, which is central to everyday social interactions. Although such reasoning is inherently multimodal, it remains largely unexplored by existing multimodal large language models (MLLMs). To address this gap, we introduce PIVOTS, the first benchmark built from Social-IQ 2.0 and YouTube data to evaluate MLLMs' ability to predict bidirectional interpersonal relationship dimensions grounded in established psychology research. In addition, PIVOTS includes auxiliary tasks that assess models' ability to identify and leverage the critical visual cues underlying such predictions. We evaluate both proprietary and open-source MLLMs and conduct detailed ablation studies to analyze the effects of visual modalities and explicit social role information in conversational utterances. We further examine how joint and pairwise prediction settings benefit MLLMs in scoring bidirectional PIVOTS dimensions. Project page and resources: https://flynnzhangsx.github.io/PIVOTSBench/ .

多模态社会推理基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。