提出新型时空递归结构变压器,提升多智能体离线强化学习泛化能力。
STAIRS-Former: Spatio-Temporal Attention with Interleaved Recursive Structure Transformer for Offline Multi-task Multi-agent Reinforcement Learning
- 设计时空交错递归结构,增强智能体间注意力与历史依赖捕捉
- 在SMAC、MPE等6个基准上超越现有方法,性能达新纪录
- 适合需要处理异构智能体数量和复杂协作的多任务场景
离线多智能体强化学习(MARL)在多任务数据集上面临挑战:不同任务中智能体数量不一,且需泛化至未见场景。现有方法虽采用观察分块与分层技能学习的Transformer架构,但对智能体间注意力利用不足,且仅依赖单一历史令牌,难以捕捉部分可观测设置下的长时序依赖。本文提出STAIRS-Former,一种融合空间与时间层次结构的Transformer架构,可有效关注关键令牌并建模长期交互历史。同时引入令牌丢弃机制以增强对不同智能体规模的鲁棒性与泛化能力。在SMAC、SMAC-v2、MPE和MaMuJoCo等多个多智能体基准上的广泛实验表明,该方法持续优于先前方法,取得新的最佳性能。
原文摘要 · Abstract (English)
Offline multi-agent reinforcement learning (MARL) with multi-task datasets is challenging due to varying numbers of agents across tasks and the need to generalize to unseen scenarios. Prior works employ transformers with observation tokenization and hierarchical skill learning to address these issues. However, they underutilize the transformer attention mechanism for inter-agent coordination and rely on a single history token, which limits their ability to capture long-horizon temporal dependencies in partially observable MARL settings. In this paper, we propose STAIRS-Former, a transformer architecture augmented with spatial and temporal hierarchies that enables effective attention over critical tokens while capturing long interaction histories. We further introduce token dropout to enhance robustness and generalization across varying agent populations. Extensive experiments on diverse multi-agent benchmarks, including SMAC, SMAC-v2, MPE, and MaMuJoCo, with multi-task datasets demonstrate that STAIRS-Former consistently outperforms prior methods and achieves new state-of-the-art performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。