arXiv:2511.11083cs.LGcs.AI2025-11AAAI被引 1

提出高效训练框架ScaPT,实现大规模群体协作的零样本协同。

Efficient Reinforcement Learning for Zero-Shot Coordination in Evolving Games

  • 用元代理机制在小模型中模拟大规模群体,节省计算资源。
  • 在汉诺比游戏中实现比现有方法更优的零样本协作表现。
  • 适合研究大规模多智能体协作与高效训练方法的学者。

零样本协同(ZSC)是多智能体博弈中的关键挑战,近年来在强化学习领域备受关注,尤其在复杂动态博弈中。其核心是让智能体具备泛化能力,无需微调即可与未见过的多样化、可能持续演化的合作者有效协作。基于种群的训练方法虽能近似这种动态合作环境并取得良好零样本协同性能,但受限于计算资源,现有方法主要聚焦于小规模种群的多样性优化,忽视了扩大种群规模带来的潜在性能提升。为此,本文提出可扩展种群训练(ScaPT)框架,包含两个关键组件:一是元代理,通过跨智能体选择性共享参数高效实现大规模种群;二是互信息正则项,确保种群多样性。为验证ScaPT的有效性,本文在汉诺比(Hanabi)合作游戏中对其及代表性框架进行了评估,结果表明其具有显著优势。

原文摘要 · Abstract (English)

Zero-shot coordination(ZSC), a key challenge in multi-agent game theory, has become a hot topic in reinforcement learning (RL) research recently, especially in complex evolving games. It focuses on the generalization ability of agents, requiring them to coordinate well with collaborators from a diverse, potentially evolving, pool of partners that are not seen before without any fine-tuning. Population-based training, which approximates such an evolving partner pool, has been proven to provide good zero-shot coordination performance; nevertheless, existing methods are limited by computational resources, mainly focusing on optimizing diversity in small populations while neglecting the potential performance gains from scaling population size. To address this issue, this paper proposes the Scalable Population Training (ScaPT), an efficient RL training framework comprising two key components: a meta-agent that efficiently realizes a population by selectively sharing parameters across agents, and a mutual information regularizer that guarantees population diversity. To empirically validate the effectiveness of ScaPT, this paper evaluates it along with representational frameworks in Hanabi cooperative game and confirms its superiority.

多智能体强化学习零样本协同高效训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。