arXiv:2509.14485cs.AI2025-09

用新方法分析智能体合作能力,发现高分不等于真合作。

Beyond the high score: Prosocial ability profiles of multi-agent populations

  • 用贝叶斯测量布局推断多智能体的合作能力画像
  • 部分低分智能体实际合作能力更强,高分未必真协作
  • 揭示竞赛评估体系漏洞,适合研究AI社会行为的学者

AI智能体社会能力的开发与评估需要复杂环境以自然催生竞争与合作行为。尽管博弈论可解释某些团队或智能体群体为何表现更优,但像遵循惯例这类抽象行为在训练与评估中更难控制。梅尔特инг杯竞赛是用于评估AI系统协作能力的社会化AI评测套件。本文应用一种名为测量布局(Measurement Layouts)的贝叶斯方法,推断梅尔特丁杯竞赛中多智能体系统的能力画像。结果表明,这些能力画像不仅能预测未来在梅尔特丁杯中的表现,还能揭示智能体的潜在亲社会能力。分析显示,虽然更高亲社会能力有时与更好表现相关,但这并非普遍规律——部分低分智能体展现出更强的协作能力。此外,顶尖竞赛提交方案更可能在无需亲社会能力的场景中取得高分。结合报告称冠军团队使用针对特定环境的硬编码解决方案,提示至少一支顶尖队伍可能优化了无需合作的条件,可能利用了评估框架的局限性。我们提出改进协作需求标注的建议,并为未来研究方向提供思路,以应对不同测试环境引入的偏差。结果表明,测量布局不仅具备强预测准确率,还提供可操作洞察,推动复杂社交环境中AI系统评估的透明化与通用化。

原文摘要 · Abstract (English)

The development and evaluation of social capabilities in AI agents require complex environments where competitive and cooperative behaviours naturally emerge. While game-theoretic properties can explain why certain teams or agent populations outperform others, more abstract behaviours, such as convention following, are harder to control in training and evaluation settings. The Melting Pot contest is a social AI evaluation suite designed to assess the cooperation capabilities of AI systems. In this paper, we apply a Bayesian approach known as Measurement Layouts to infer the capability profiles of multi-agent systems in the Melting Pot contest. We show that these capability profiles not only predict future performance within the Melting Pot suite but also reveal the underlying prosocial abilities of agents. Our analysis indicates that while higher prosocial capabilities sometimes correlate with better performance, this is not a universal trend-some lower-scoring agents exhibit stronger cooperation abilities. Furthermore, we find that top-performing contest submissions are more likely to achieve high scores in scenarios where prosocial capabilities are not required. These findings, together with reports that the contest winner used a hard-coded solution tailored to specific environments, suggest that at least one top-performing team may have optimised for conditions where cooperation was not necessary, potentially exploiting limitations in the evaluation framework. We provide recommendations for improving the annotation of cooperation demands and propose future research directions to account for biases introduced by different testing environments. Our results demonstrate that Measurement Layouts offer both strong predictive accuracy and actionable insights, contributing to a more transparent and generalisable approach to evaluating AI systems in complex social settings.

多智能体合作评估社会智能测评框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。