破解多模态大模型的思维之谜,让机器像人一样理解他人想法。
From Black Boxes to Transparent Minds: Evaluating and Enhancing the Theory of Mind in Multimodal Large Language Models
- 通过分析注意力头机制,揭示多模态模型如何区分不同视角的认知信息。
- 在自建数据集GridToM上,模型展现出可解释的信念推理能力。
- 无需训练,仅调整注意力方向即可显著提升模型的共情推理表现。
随着大语言模型的发展,人们期待它们能具备类人的心智理论(ToM)以辅助日常任务。然而,现有机器心智理论评估方法主要针对单模态模型,且多将其视为黑箱,缺乏对内部机制的可解释性探究。为此,本研究基于内部机制提出一种可解释性评估方法,用于分析多模态大语言模型(MLLMs)的心智理论能力。我们首先构建了一个多模态心智理论测试数据集GridToM,包含多种信念测试任务及多视角感知信息。分析表明,多模态大模型中的注意力头能够区分不同视角的认知信息,为心智理论能力提供证据。此外,我们提出一种轻量级、无需训练的方法,通过沿注意力头方向进行调整,显著增强模型表现出的心智理论能力。
原文摘要 · Abstract (English)
As large language models evolve, there is growing anticipation that they will emulate human-like Theory of Mind (ToM) to assist with routine tasks. However, existing methods for evaluating machine ToM focus primarily on unimodal models and largely treat these models as black boxes, lacking an interpretative exploration of their internal mechanisms. In response, this study adopts an approach based on internal mechanisms to provide an interpretability-driven assessment of ToM in multimodal large language models (MLLMs). Specifically, we first construct a multimodal ToM test dataset, GridToM, which incorporates diverse belief testing tasks and perceptual information from multiple perspectives. Next, our analysis shows that attention heads in multimodal large models can distinguish cognitive information across perspectives, providing evidence of ToM capabilities. Furthermore, we present a lightweight, training-free approach that significantly enhances the model's exhibited ToM by adjusting in the direction of the attention head.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。