arXiv:2605.29512cs.AI2026-05被引 3

构建多智能体博弈评估平台,测试大模型的社会与策略推理能力。

MINDGAMES: A Live Arena for Evaluating Social and Strategic Reasoning in Multi-Agent LLMs

论文配图:MINDGAMES: A Live Arena for Evaluating Social and Strategic Reasoning in Multi-Agent LLMs
图 1 · 摘自论文原文
  • 设计四个游戏环境,评估信念推断、对手建模等理论心理能力。
  • 944个智能体参与竞赛,发现顶尖系统依赖结构化设计而非真正推理。
  • 释放2.9万场对战数据,提供可复现的离线评测协议。

大型语言模型正被广泛用作交互式代理,但其在长期多智能体互动中进行社会与策略推理的能力仍不清晰。现有评估依赖静态情景或单轮游戏基准,难以捕捉现实多智能体场景所需的持续、多维度推理。我们提出Mindgames,一个基于TextArena的多游戏竞技场与评估平台,涵盖信念推断(隐藏信息下)、对手建模(重复博弈)、合作推理(知识不对称)及持续欺骗(社交推理)等核心能力。平台提供统一交互界面、TrueSkill评分系统及全程轨迹记录,通过2025年某重大人工智能会议的竞赛周期,评估了来自76个团队的944个智能体在四个游戏(Colonel Blotto、迭代囚徒困境、猜词游戏、秘密玛菲亚)中的表现。分析揭示:智能体普遍存在规则僵化问题,顶级系统高度依赖显式结构支撑,且排行榜有效性在不同环境中差异显著。尤其在失败密集型环境如秘密玛菲亚中,系统对对手错误的鲁棒性常被误认为战略能力。我们公开发布包含29,571场多智能体对战的全轨迹数据集,并推出MG-Ref协议,可在固定参考池上对新智能体进行确定性离线评测,采用与本研究一致的错误归因标准。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly deployed as interactive agents, yet their capacity for social and strategic reasoning over extended interaction remains poorly understood. Existing evaluations rely on static vignettes or single-game benchmarks that cannot capture the sustained, multi-faceted reasoning that real-world multi-agent settings demand. We introduce Mindgames, a multi-game arena and evaluation platform for LLM agents that operationalizes complementary reasoning demands relevant to ``theory of mind'': belief attribution under hidden information, opponent modeling through repeated strategic interaction, cooperative inference under knowledge asymmetries, and sustained deception in social deduction. Built on TextArena, Mindgames provides a unified interaction interface, TrueSkill-based rating, and full trajectory logging across four game environments. We instantiate Mindgames through a 2025 competition cycle hosted at a major AI conference, which assessed 944 submitted agents from 76 teams across four games: Colonel Blotto, Iterated Prisoner's Dilemma, Codenames, and Secret Mafia. Our analysis surfaces both agent-level and evaluation-level limitations: brittle rule adherence remains a major bottleneck, top-performing systems repeatedly rely on explicit structural scaffolding, and leaderboard validity differs sharply across environments. In particular, failure-heavy environments can reward robustness to opponent errors as much as strategic ability, with Secret Mafia exhibiting a pronounced error-survival confound in this cycle. We release a dataset of 29,571 multi-agent games with turn-level observations, actions, and rewards, together with MG-Ref, a deterministic offline tournament protocol that scores new agents against a frozen reference pool of top-ranked, low-error Stage~II submissions under the same error-attribution lens used in this analysis.

多智能体推理评估博弈测试大模型评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。