arXiv:2605.04312cs.AIcs.MA2026-05

用多智能体博弈构建抗饱和与污染的动态评测基准

Agent Island: A Saturation- and Contamination-Resistant Benchmark from Multiagent Games

论文配图:Agent Island: A Saturation- and Contamination-Resistant Benchmark from Multiagent Games
图 1 · 摘自论文原文
  • 设计多智能体对抗环境,让语言模型在合作与说服中竞争
  • GPT-5.5以5.64分领先,显著优于第二名(3.10)和第三名(2.86)
  • 可追踪模型进步,适合评估大模型在动态协作中的真实能力

静态能力评测存在饱和与污染问题,难以持续跟踪进展。我们提出Agent Island,一个多人模拟环境,语言模型在此进行合作、冲突与说服博弈。该环境构建了动态基准,新模型总能超越当前最强者,且对手为不断进化的自适应智能体。采用贝叶斯Plackett-Luce模型对玩家排名,量化技能不确定性。在999场涉及49个不同模型的比赛中,openai/gpt-5.5以5.64的后验均值技能领先,第二名openai/gpt-5.2为3.10,第三名openai/gpt-5.3-codex为2.86。我们发布游戏日志数据集,用于分析模型行为。例如,发现同厂商模型在决赛投票中更倾向支持本方,偏好度高出8.3个百分点;此效应在OpenAI模型中最强,Anthropic最弱。

原文摘要 · Abstract (English)

Static capabilities benchmarks suffer from saturation and contamination, making it difficult to track capabilities progress over time. We introduce Agent Island, a multiplayer simulation environment in which language-model agents compete in a game of interagent cooperation, conflict, and persuasion. The environment yields a dynamic benchmark designed to mitigate both saturation and contamination; new models can always outperform the current leading player in this winner-take-all game, and agents compete against other adaptive agents rather than face a fixed task set. We rank players with a Bayesian Plackett-Luce model, allowing us to quantify uncertainty in player skill. In 999 games involving 49 unique models, openai/gpt-5.5 dominates its peers with a posterior mean skill of 5.64, compared with 3.10 for the second-ranked model, openai/gpt-5.2, and 2.86 for the third-ranked model, openai/gpt-5.3-codex. We release the game logs as a dataset for analyses of model behavior. As an example, we investigate same-provider preference in final-round votes and find that models are 8.3 p.p. more likely to support a same-provider finalist than finalists from other providers. This preference is not uniform across providers: among separately estimated providers, the effect is strongest for OpenAI models and weakest for Anthropic models.

大模型评测多智能体动态基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。