arXiv:2609.03667cs.LGcs.AI2026-09

提升离线多智能体强化学习泛化能力,关键在训练任务多样性。

Out-of-Distribution Generalisation with Sequence Models in Offline Multi-Agent Reinforcement Learning

论文配图:Out-of-Distribution Generalisation with Sequence Models in Offline Multi-Agent Reinforcement Learning
图 1 · 摘自论文原文
  • 扩展序列模型支持多任务与可变智能体数量
  • 任务多样性比数据量对零样本泛化影响更大
  • 四类环境测试中性能提升3.2倍,优于行为克隆

离线多智能体强化学习中的未知任务泛化仍是核心挑战。本文对离线设置下的零样本任务泛化进行了系统分析,并深入探究了任务多样性、数据集规模与网络容量之间的缩放规律。为支持该研究,我们扩展了离线序列建模架构,以处理多任务观测与动作空间以及任务间可变的智能体数量。主要发现表明,扩大任务多样性(而非单纯增加数据量)是实现稳健零样本迁移的关键因素。在四个挑战性环境(Connector、RWARE、SMAX、LBF)上的大规模实验显示,所提多任务方法在未见测试任务上平均性能提升3.2倍,且持续优于强基线行为克隆模型。结果表明,构建可泛化的多智能体强化学习代理应优先关注训练分布的任务多样性,尤其包含不同智能体数量的情形,为有效扩展离线多智能体强化学习提供了清晰路径。

原文摘要 · Abstract (English)

Generalising to unseen tasks remains a fundamental challenge in offline multi-agent reinforcement learning (MARL). In this work, we present a principled analysis of zero-shot task generalisation in the offline setting and conduct an extensive empirical investigation into the scaling behaviour governing task diversity, dataset size, and network capacity. To facilitate this study, we extend offline sequence modelling architectures to handle multi-task observation and action spaces alongside variable agent counts across tasks. Our primary finding is that scaling task diversity---rather than sheer dataset size is the dominant factor in achieving robust zero-shot transfer. Through large-scale experiments across four challenging environments (Connector, RWARE, SMAX, and LBF), we demonstrate that our multi-task approach achieves a mean improvement of 3.2x on held-out test tasks compared to single-task models and consistently outperforms strong behaviour cloning baselines. These results suggest that the development of generalisable MARL agents should prioritise the diversity of the training distribution with varying numbers of agents, providing a roadmap for scaling offline MARL effectively.

多智能体离线强化学习泛化能力序列建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。