评测多模态大模型在多人会议中理解他人心理状态的能力
MeetingToM: Evaluating Multimodal LLMs on Theory-of-Mind Reasoning in Multi-Party Meetings

- 构建分层级的会议心理推理基准,涵盖个体、两人和群体三个层面
- 发现模型难以区分表面共识与私下异议,尤其在社会压力下
- 适合研究社交智能、人机交互与多模态推理的学者使用
心智理论(ToM)是理解他人信念、意图和知识状态的核心能力,对社交互动至关重要,但当前多模态大语言模型(MLLMs)在多人会议场景中仍面临挑战,因线索分散于语音与行为之中。现有多模态ToM基准主要关注基于视频的可验证信号问答,对隐性社会状态与群体动态覆盖有限。本文提出MeetingToM,一个面向自然对话中复杂社会行为推理的基准,聚焦会议特有现象如‘伪共识’——表面一致实则隐藏分歧。该基准分层设计,评估三类推理能力:(i) 个体心理状态预测,(ii) 二人间交流对象理解,(iii) 群体共识推理。提供统一评估协议,系统分析代表性MLLMs,揭示其在融合非语言线索、推断隐藏态度及辨别真实共识与伪共识方面存在持续局限。结果凸显关键挑战,并确立MeetingToM作为推动会议场景下多模态心智理论发展的测试平台。
原文摘要 · Abstract (English)
Theory of Mind (ToM), the ability to infer other's beliefs, intentions, and states of knowledge, is central to social interaction, yet remains challenging for current Multimodal Large Language Models (MLLMs), especially in multi-party meetings where cues are distributed across speech and behavior. Existing multimodal ToM benchmarks mainly focus on video-grounded question answering over overt, externally verifiable signals, and provide limited coverage of latent social states and group dynamics. We introduce MeetingToM, a benchmark for complex social behavior reasoning in naturalistic multi-party meetings. MeetingToM targets meeting-specific phenomena such as \textbf{pseudo-consensus}, where apparent agreement masks private dissent under social pressure. The benchmark is hierarchically organized to evaluate ToM at increasing levels of social granularity, including (i) subject-level mental state prediction, (ii) dyadic-level addressee understanding, and (iii) group-level consensus reasoning. We provide a unified evaluation protocol and conduct systematic analyses of representative MLLMs, revealing persistent limitations in integrating non-verbal cues, inferring hidden attitudes, and distinguishing genuine consensus from pseudo-consensus. Our results highlight key challenges and establish MeetingToM as a testbed for advancing meeting-grounded ToM in multimodal models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。