用游戏化评测大模型多智能体推理与心智理论能力,发现顶尖模型表现反而更差。
The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind
- 设计游戏化基准Decrypto,剥离干扰因素专注评测心智理论能力
- 顶尖模型在经典认知实验变体中表现逊于旧模型和基础词向量方法
- 首个支持人机交互的可复现心智理论实验平台,适合评估智能体协作与推理
随着大语言模型具备代理能力,需在复杂多智能体场景中与人类及其他智能体合作或竞争,这要求具备心智理论(ToM)——即推理其他智能体‘心理状态’的能力。然而,现有评测基准存在范围狭窄、数据泄露、性能饱和及缺乏交互性等问题,导致对LLM的多智能体能力理解不足。为此,我们提出Decrypto,一个受认知科学、计算语用学和多智能体强化学习启发的游戏化基准。其设计力求在其他维度尽可能简单,消除常见混淆因素。据我们所知,这也是首个支持交互式心智理论实验的设计平台。通过前沿LLM的全面实证评估、鲁棒性测试及人机跨玩实验验证了基准有效性。结果显示,当前大模型的游戏能力远低于人类,甚至不如简单的词嵌入基线。我们在Decrypto中构建了两个经典认知科学实验的变体,用于评估三种关键的心智理论能力。令人惊讶的是,最先进推理模型在这些任务上的表现显著劣于较老模型。这表明Decrypto填补了当前推理与心智理论评估的关键空白,并为构建更优人工智能体指明方向。
原文摘要 · Abstract (English)
As Large Language Models (LLMs) gain agentic abilities, they will have to navigate complex multi-agent scenarios, interacting with human users and other agents in cooperative and competitive settings. This will require new reasoning skills, chief amongst them being theory of mind (ToM), or the ability to reason about the "mental" states of other agents. However, ToM and other multi-agent abilities in LLMs are poorly understood, since existing benchmarks suffer from narrow scope, data leakage, saturation, and lack of interactivity. We thus propose Decrypto, a game-based benchmark for multi-agent reasoning and ToM drawing inspiration from cognitive science, computational pragmatics and multi-agent reinforcement learning. It is designed to be as easy as possible in all other dimensions, eliminating confounding factors commonly found in other benchmarks. To our knowledge, it is also the first platform for designing interactive ToM experiments. We validate the benchmark design through comprehensive empirical evaluations of frontier LLMs, robustness studies, and human-AI cross-play experiments. We find that LLM game-playing abilities lag behind humans and simple word-embedding baselines. We then create variants of two classic cognitive science experiments within Decrypto to evaluate three key ToM abilities. Surprisingly, we find that state-of-the-art reasoning models are significantly worse at those tasks than their older counterparts. This demonstrates that Decrypto addresses a crucial gap in current reasoning and ToM evaluations, and paves the path towards better artificial agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。