提出新方法解决离线多智能体强化学习中的分布偏移问题。
Offline Multi-agent Reinforcement Learning via Sequential Score Decomposition
- 通过序列化得分分解,从联合行为策略中提取个体正则化信号。
- 在多个粒子环境和Multi-agent MuJoCo上达到当前最优性能。
- 适合研究离线多智能体学习与策略泛化方向的学者。
离线协作式多智能体强化学习因联合动作空间高维及分布外动作选择导致分布偏移,面临独特挑战。本文指出,合作任务的多均衡特性引发高度多模态的联合行为策略空间与质量不一的行为数据,使个体策略正则化难以对齐一致协调模式,从而引发策略分布偏移。为此,我们设计了一种序列化得分函数分解方法,从联合行为策略中提炼出各智能体的正则化信号,在去中心化执行约束下引导协调模态选择。随后利用灵活的基于扩散的生成模型,从多模态离线数据中学习这些得分函数,并将其集成至联合动作评价器中,指导策略更新向高奖励、分布内区域收敛。该方法在多个粒子环境与Multi-agent MuJoCo基准上表现持续领先。据我们所知,这是首个明确解决离线与在线MARL间分布差距的工作,为更通用的基于策略的离线多智能体学习方法开辟了道路。
原文摘要 · Abstract (English)
Offline cooperative multi-agent reinforcement learning (MARL) faces unique challenges due to distributional shifts, particularly stemming from the high dimensionality of joint action spaces and the presence of out-of-distribution joint action selections. In this work, we highlight that a fundamental challenge in offline MARL arises from the multi-equilibrium nature of cooperative tasks, which induces a highly multimodal joint behavior policy space coupled with heterogeneous-quality behavior data. This makes it difficult for individual policy regularization to align with a consistent coordination pattern, leading to the policy distribution shift problems. To tackle this challenge, we design a sequential score function decomposition method that distills per-agent regularization signals from the joint behavior policy, which induces coordinated modality selection under decentralized execution constraints. Then we leverage a flexible diffusion-based generative model to learn these score functions from multimodal offline data, and integrate them into joint-action critics to guide policy updates toward high-reward, in-distribution regions under a shared team reward. Our approach achieves state-of-the-art performance across multiple particle environments and Multi-agent MuJoCo benchmarks consistently. To the best of our knowledge, this is the first work to explicitly address the distributional gap between offline and online MARL, paving the way for more generalizable offline policy-based MARL methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。