首个基于心理学的多模态心智理论评测基准,评估模型理解他人心理状态能力。
CoMMET: A Psychologically Grounded Benchmark for Evaluating Theory of Mind in Multimodal LLMs
- 构建多模态、多轮、开放式的心理状态评测数据集
- 覆盖多种心理状态,首次实现跨模态心智推理综合评估
- 适合研究多模态大模型社会智能与人类认知对齐的研究者
心智理论(ToM)——即对自身及他人心理状态进行推理的能力——是人类社会智能的核心。随着多模态大语言模型(MLLMs)在真实场景中的广泛应用,验证其是否具备此类社会推理能力至关重要。然而,现有评估基准多局限于文本输入,且仅关注信念类任务。本文提出全新多模态基准数据集CoMMET,灵感来自心智理论小册子任务,涵盖更广泛的心理状态并引入多轮交互测试。据我们所知,这是首个基于心理学、在多模态、开放式、多轮情境下评估MLLMs在多种心理状态上的基准。通过全面评估不同家族和规模的模型,分析了当前模型的优劣,并指明未来改进方向。
原文摘要 · Abstract (English)
Theory of Mind (ToM)-the ability to reason about the mental states of oneself and others-is a cornerstone of human social intelligence. As Multimodal Large Language Models (MLLMs) become ubiquitous in real-world applications, validating their capacity for this level of social reasoning is essential for effective and natural interactions. However, existing benchmarks for assessing ToM in MLLMs are limited; most rely solely on text inputs and focus narrowly on belief-related tasks. In this paper, we propose a new multimodal benchmark dataset, CoMMET, a comprehensive mental states and moral evaluation task inspired by the Theory of Mind Booklet Task. CoMMET expands the scope of evaluation by covering a broader range of mental states and introducing multi-turn testing. To the best of our knowledge, this is the first psychology-grounded benchmark to evaluate MLLMs across multiple mental states in a multimodal, open-ended, and multi-turn setting. Through a comprehensive assessment of MLLMs across different families and sizes, we analyze the strengths and limitations of current models and identify directions for future improvement.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。