让智能体学会根据角色调整策略,应对未知对手。
Role Play: Learning Adaptive Role-Specific Strategies in Multi-Agent Interactions
- 用角色嵌入替代传统策略池,实现策略多样性
- 在厨艺、采集等任务中超越基线模型表现
- 适合需要动态协作与竞争的复杂多智能体场景
多智能体强化学习中的零样本协调问题日益受到关注,要求智能体能适应未见过的对手。传统自对弈框架虽能生成多样策略,但在需平衡合作与竞争的真实场景中仍存在局限。受社会价值取向启发,本文提出新框架「角色扮演」(Role Play, RP),通过角色嵌入将策略多样性转化为角色多样性。该框架训练统一策略并结合角色预测器估计其他智能体的联合角色嵌入,使自身可动态适应所处角色。理论证明,优化相对于近似角色策略的期望累积奖励可逼近最优策略。在协作型(Overcooked)和混合动机游戏(Harvest, CleanUp)中的实验表明,RP在与未知智能体交互时持续优于强基线,展现了出色的鲁棒性与适应能力。
原文摘要 · Abstract (English)
Zero-shot coordination problem in multi-agent reinforcement learning (MARL), which requires agents to adapt to unseen agents, has attracted increasing attention. Traditional approaches often rely on the Self-Play (SP) framework to generate a diverse set of policies in a policy pool, which serves to improve the generalization capability of the final agent. However, these frameworks may struggle to capture the full spectrum of potential strategies, especially in real-world scenarios that demand agents balance cooperation with competition. In such settings, agents need strategies that can adapt to varying and often conflicting goals. Drawing inspiration from Social Value Orientation (SVO)-where individuals maintain stable value orientations during interactions with others-we propose a novel framework called \emph{Role Play} (RP). RP employs role embeddings to transform the challenge of policy diversity into a more manageable diversity of roles. It trains a common policy with role embedding observations and employs a role predictor to estimate the joint role embeddings of other agents, helping the learning agent adapt to its assigned role. We theoretically prove that an approximate optimal policy can be achieved by optimizing the expected cumulative reward relative to an approximate role-based policy. Experimental results in both cooperative (Overcooked) and mixed-motive games (Harvest, CleanUp) reveal that RP consistently outperforms strong baselines when interacting with unseen agents, highlighting its robustness and adaptability in complex environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。