打造超大规模多智能体强化学习新基准,挑战长期协作与泛化能力。
Multi-Agent Craftax: Benchmarking Open-Ended Multi-Agent Reinforcement Learning at the Hyperscale
- 基于JAX的Craftax扩展,支持多智能体并行训练,2.5亿步训练<1小时
- 引入异构智能体与交易机制,需复杂协作才能达成目标
- 揭示现有算法在长程信用分配与探索上的短板,适合前沿MARL研究
多智能体强化学习(MARL)的发展需要能真正考验当前方法极限的挑战性基准。然而,现有基准多聚焦于短周期、狭窄任务,难以充分评估多智能体系统中固有的长期依赖与泛化能力。为此,我们首先提出Craftax-MA:对流行开源开放世界强化学习环境Craftax的扩展,支持多智能体,并在一个环境中评估多种通用能力。该环境基于JAX实现,运行极快,使用2.5亿次环境交互的训练任务可在一小时内完成。为进一步提升挑战性,我们还推出了Craftax-Coop,引入异构智能体、交易等机制,要求智能体间进行复杂协作方能成功。分析表明,现有算法在长时程信用分配、探索与协作等方面表现不佳,凸显该基准推动MARL长期研究的潜力。
原文摘要 · Abstract (English)
Progress in multi-agent reinforcement learning (MARL) requires challenging benchmarks that assess the limits of current methods. However, existing benchmarks often target narrow short-horizon challenges that do not adequately stress the long-term dependencies and generalization capabilities inherent in many multi-agent systems. To address this, we first present \textit{Craftax-MA}: an extension of the popular open-ended RL environment, Craftax, that supports multiple agents and evaluates a wide range of general abilities within a single environment. Written in JAX, \textit{Craftax-MA} is exceptionally fast with a training run using 250 million environment interactions completing in under an hour. To provide a more compelling challenge for MARL, we also present \textit{Craftax-Coop}, an extension introducing heterogeneous agents, trading and more mechanics that require complex cooperation among agents for success. We provide analysis demonstrating that existing algorithms struggle with key challenges in this benchmark, including long-horizon credit assignment, exploration and cooperation, and argue for its potential to drive long-term research in MARL.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。