arXiv:2605.02063cs.MAcs.AI2026-05

首个支持战略共谋的多智能体强化学习基准平台,可模拟真实合作竞争关系。

Coopetition-Gym v1: A Formally Grounded Platform for Mixed-Motive Multi-Agent Reinforcement Learning under Strategic Coopetition

  • 构建20个环境,分离收益与奖励,支持三种奖励模式切换。
  • 4个历史案例复现准确率达81.7%~98.3%,验证平台可信度。
  • 提供126种基线算法,适配研究博弈、协作与竞争的学者使用。

我们提出Coopetition-Gym v1,一个面向策略性共谋下的混合动机多智能体强化学习的基准平台。平台包含20个环境,分为四类机制:相互依存与互补性(arXiv:2510.18802)、信任与声誉动态(arXiv:2510.24909)、集体行动与忠诚度(arXiv:2601.16237)、顺序交互与互惠性(arXiv:2604.01240)。每个环境具有闭式收益结构和基于对应报告校准的相互依存矩阵。奖励层可配置为私有、集成、合作三种结构模式,实现奖励类型消融分析。四个环境基于历史共谋关系校准,对三星-索尼液晶屏、雷诺-日产联盟、Apache HTTP Server、苹果iOS应用商店的复现准确率分别为98.3%、81.7%、86.7%、87.3%。平台支持Gymnasium、PettingZoo Parallel和PettingZoo AEC接口,附带126个参考算法:16种学习算法、7种博弈论基准、2种启发式基线、101种恒定动作策略。一项参考实验在所有环境与奖励配置下训练16种学习算法,共25,708次训练运行,生成1,116次行为审计数据,均以CC-BY-4.0许可发布,并采用Croissant 1.0元数据标准。Coopetition-Gym v1是首个融合连续动作、参数化奖励互惠性、校准相互依存系数、博弈论基准与验证案例的平台。

原文摘要 · Abstract (English)

We present Coopetition-Gym v1, a benchmark platform for mixed-motive multi-agent reinforcement learning under strategic coopetition. The platform comprises twenty environments organized into four mechanism classes that correspond to four foundational technical reports: interdependence and complementarity (arXiv:2510.18802), trust and reputation dynamics (arXiv:2510.24909), collective action and loyalty (arXiv:2601.16237), and sequential interaction and reciprocity (arXiv:2604.01240). Each environment carries a closed-form payoff structure and a calibrated interdependence matrix derived from the corresponding report. Every environment exposes a parameterized reward layer configurable across three structurally distinct modes (private, integrated, cooperative). This separation of payoff from reward enables reward-type ablation, the platform's principal methodological apparatus. Four of the twenty environments are calibrated against historically documented coopetitive relationships and reproduce their outcomes at 98.3, 81.7, 86.7, and 87.3 percent on the validation rubric (Samsung-Sony LCD, Renault-Nissan Alliance, Apache HTTP Server, Apple iOS App Store). The platform exposes Gymnasium, PettingZoo Parallel, and PettingZoo AEC interfaces and ships 126 reference algorithms: 16 learning algorithms, 7 game-theoretic oracles, 2 heuristic baselines, and 101 constant-action policies. A reference experimental study trained the 16 learning algorithms on every environment under every reward configuration with seven random seeds, producing a 25,708-run training corpus and a 1,116-run behavioral audit corpus, both released under CC-BY-4.0 with Croissant 1.0 metadata. Coopetition-Gym v1 is the first platform to combine continuous-action mixed-motive environments, parameterized reward mutuality, calibrated interdependence coefficients, game-theoretic oracle baselines, and validated case studies.

多智能体强化学习共谋博弈基准平台

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。