arXiv:2508.12524cs.LG2025-08

NeurIPS 2023 多任务强化学习竞赛展示通用策略在新任务上的卓越表现。

Results of the NeurIPS 2023 Neural MMO Competition on Multi-task Reinforcement Learning

  • 训练目标条件策略,实现跨任务、跨地图、跨对手的泛化能力。
  • 顶尖方案在单张4090显卡上8小时内得分达基线4倍。
  • 开源全部代码与模型权重,支持复现与后续研究。

我们公布了 NeurIPS 2023 Neural MMO 竞赛的结果,吸引了超过200名参与者和提交作品。参赛者训练了可在未见任务、地图和对手上泛化的目标条件策略。顶级方案在单张4090显卡上仅用8小时训练,得分达到基线的4倍。我们已将 Neural MMO 及竞赛相关所有内容以 MIT 许可证开源,包括基线和顶尖方案的策略权重与训练代码。

原文摘要 · Abstract (English)

We present the results of the NeurIPS 2023 Neural MMO Competition, which attracted over 200 participants and submissions. Participants trained goal-conditional policies that generalize to tasks, maps, and opponents never seen during training. The top solution achieved a score 4x higher than our baseline within 8 hours of training on a single 4090 GPU. We open-source everything relating to Neural MMO and the competition under the MIT license, including the policy weights and training code for our baseline and for the top submissions.

强化学习多任务竞赛结果

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。