首个评测多模态大模型群体心智能力的基准,揭示其在社会非线性互动中的不足。
GroupToM-Bench: Benchmarking Group Theory of Mind and Nonlinear Social Emergence in MLLMs

- 构建七级认知审计框架,贯通个体信念到群体结果的因果链
- 模型在群体行为预测上显著落后于人类,尤其无法捕捉非线性社会动态
- 适合研究社会智能、群体行为建模与多智能体系统的研究者
真正通用智能不仅需要对物理世界的理解,还需对社会世界的建模:即推断个体心理状态如何相互作用并凝聚为群体结果。尽管个体层面的心智理论(ToM)推理取得进展,现有多模态大语言模型仍无法完成这一更复杂任务。集体行为源于社会张力、从众机制与结构约束的非线性交互,无法通过简单叠加个体意图还原。我们提出GroupToM-Bench,首个面向群体层面心智理论的多模态基准,围绕从微观层面对象的BDI状态(信念、欲望、意图),中观层面群体张力与结构约束,到宏观层面结果预测与机制归因的完整因果链构建。为探测这一全链条,我们开发了七级认知审计框架。实验表明当前模型与人类基线存在显著差距,暴露出对社会结构和非线性集体动态处理能力的缺失。
原文摘要 · Abstract (English)
True general intelligence requires not only a model of the physical world but also a social world model: the capacity to infer how individual mental states interact and crystallize into group-level outcomes. Despite notable progress in individual-level Theory of Mind (ToM) reasoning, existing multimodal large language models fail at this broader task. Collective behavior emerges non-linearly from social tensions, conformity dynamics, and structural constraints, meaning it cannot be recovered by merely summing individual intentions. We present GroupToM-Bench, the first multimodal benchmark for group-level ToM, built around a causal chain spanning micro-level BDI states (belief, desire, intention), meso-level group tension and structural constraints, and macro-level outcome prediction and mechanistic attribution. To probe this full arc, we develop a seven-level cognitive audit framework. Experiments reveal a gap between current models and human baselines, highlighting a failure to process social structures and non-linear collective dynamics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。