通过动态组合离线与在线动作,实现高效多智能体强化学习迁移。
Sim2O: Efficient Offline-to-Online MARL via Joint Action Composition

- 将多智能体策略迁移分解为动作的动态组合过程。
- 在多个基准上性能显著优于现有方法,无需额外训练目标。
- 适合需要低成本协同决策的复杂多智能体系统研究者。
离线到在线适应通过利用离线数据启动强化学习,有效缓解了在线探索的高昂成本。尽管该范式在单智能体场景中已有广泛研究,但在多智能体强化学习(MARL)中的拓展仍基本空白,而后者对复杂协同决策至关重要。为此,我们提出 Sim2O,一个简洁、极简的离线到在线 MARL 框架。不同于将适应视为单一联合决策,Sim2O 将其建模为可组合的过程:通过在各智能体间动态融合离线与在线动作提议,合成候选联合动作。借助中心化价值函数评估这些混合组合,Sim2O 能够识别高价值协调策略,且无需辅助训练目标或结构开销。在多种基准上的实证评估表明,Sim2O 显著优于现有基线,证实了极简设计在多智能体离线到在线适应中的可行性和高效性。
原文摘要 · Abstract (English)
Offline-to-online adaptation serves as a pivotal paradigm for mitigating the prohibitive cost of online exploration by bootstrapping reinforcement learning from offline datasets. While this paradigm has been extensively studied in single-agent settings, its extension to Multi-Agent Reinforcement Learning (MARL) remains largely unexplored, despite its critical relevance to complex coordinated decision-making. To bridge this gap, we introduce Sim2O, an elegant and minimalist framework for offline-to-online MARL. Rather than treating adaptation as a monolithic joint decision, Sim2O conceptualizes it as a compositional process. Specifically, candidate joint actions are synthesized by dynamically blending offline and online action proposals across agents. By leveraging a centralized value function to evaluate these hybrid combinations, Sim2O identifies high-value coordination strategies without requiring auxiliary training objectives or structural overhead. Empirical evaluations across diverse benchmarks demonstrate that Sim2O significantly outperforms existing baselines, underscoring that a minimalist design is not only viable but highly effective for multi-agent offline-to-online adaptation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。