Oryx提升多智能体离线强化学习的长程协作能力
Oryx: a Scalable Sequence Model for Many-Agent Coordination in Offline MARL
- 基于保留机制与序列隐约束Q学习,设计自回归策略更新方案
- 在65个数据集上超80%任务达到顶尖性能,长轨迹保持时序一致性
- 适合研究大规模多智能体协同与离线强化学习的学者使用
离线多智能体强化学习(MARL)的核心挑战在于复杂环境中实现多智能体的长程多步协作。本文提出Oryx,一种新型离线合作式MARL算法,直接应对该挑战。Oryx借鉴近期提出的基于保留机制的Sable架构,并结合序列化的隐约束Q学习(ICQ),构建了一种新的离线自回归策略更新机制,使模型能在复杂协调任务中保持长轨迹的时序一致性。我们在多个基准测试上评估Oryx,涵盖SMAC、RWARE和Multi-Agent MuJoCo,涉及离散与连续控制任务,规模与难度各异。Oryx在65个测试数据集中的超过80%上取得当前最优表现,显著优于以往离线MARL方法,并展现出跨领域强泛化能力,尤其在多智能体与长时域场景下具备卓越可扩展性。最后,我们引入新数据集以推动离线MARL中多智能体协作的极限,进一步验证Oryx在大规模协作场景下的优异扩展能力。
原文摘要 · Abstract (English)
A key challenge in offline multi-agent reinforcement learning (MARL) is achieving effective many-agent multi-step coordination in complex environments. In this work, we propose Oryx, a novel algorithm for offline cooperative MARL to directly address this challenge. Oryx adapts the recently proposed retention-based architecture Sable and combines it with a sequential form of implicit constraint Q-learning (ICQ), to develop a novel offline autoregressive policy update scheme. This allows Oryx to solve complex coordination challenges while maintaining temporal coherence over long trajectories. We evaluate Oryx across a diverse set of benchmarks from prior works -- SMAC, RWARE, and Multi-Agent MuJoCo -- covering tasks of both discrete and continuous control, varying in scale and difficulty. Oryx achieves state-of-the-art performance on more than 80% of the 65 tested datasets, outperforming prior offline MARL methods and demonstrating robust generalisation across domains with many agents and long horizons. Finally, we introduce new datasets to push the limits of many-agent coordination in offline MARL, and demonstrate Oryx's superior ability to scale effectively in such settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。