arXiv:2506.18537cs.LGcs.MA2025-06

用Transformer建模多智能体行为,少样本高效协作

Transformer World Model for Sample Efficient Multi-Agent Reinforcement Learning

  • 基于Transformer构建世界模型,支持部分可观测下的协作预测
  • 仅5万步交互即达近优性能,显著提升样本效率
  • 适合需要高效协同的复杂多智能体任务研究者

我们提出多智能体Transformer世界模型(MATWM),一种适用于向量与图像环境的新型Transformer世界模型。MATWM结合去中心化想象框架、半中心化评判器和队友行为预测模块,使智能体在部分可观测条件下能够建模并预判他人行为。为应对非平稳性,引入优先级回放机制,让世界模型基于近期经验训练,从而适应智能体策略的动态变化。我们在星舰多智能体挑战赛、PettingZoo与MeltingPot等多个基准上评估了MATWM,结果表明其性能超越现有模型自由及先前世界模型方法,在样本效率方面表现突出,仅需约50,000次环境交互即可接近最优表现。消融实验验证了各组件的有效性,尤其在需要强协作的任务中提升显著。

原文摘要 · Abstract (English)

We present the Multi-Agent Transformer World Model (MATWM), a novel transformer-based world model designed for multi-agent reinforcement learning in both vector- and image-based environments. MATWM combines a decentralized imagination framework with a semi-centralized critic and a teammate prediction module, enabling agents to model and anticipate the behavior of others under partial observability. To address non-stationarity, we incorporate a prioritized replay mechanism that trains the world model on recent experiences, allowing it to adapt to agents' evolving policies. We evaluated MATWM on a broad suite of benchmarks, including the StarCraft Multi-Agent Challenge, PettingZoo, and MeltingPot. MATWM achieves state-of-the-art performance, outperforming both model-free and prior world model approaches, while demonstrating strong sample efficiency, achieving near-optimal performance in as few as 50K environment interactions. Ablation studies confirm the impact of each component, with substantial gains in coordination-heavy tasks.

多智能体Transformer世界模型样本效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。