arXiv:2604.08174cs.LG2026-04

用价值引导的流模型实现高效多智能体离线强化学习

Value-Guidance MeanFlow for Offline Multi-Agent Reinforcement Learning

  • 用全局优势值指导协作,将策略学习转为条件行为克隆
  • 无需调参的无分类器引导流模型,训练与推理更高效
  • 在离散和连续动作空间上表现媲美顶尖方法

离线多智能体强化学习旨在从预收集数据集中学习最优联合策略,需在最大化全局回报与缓解离线数据分布偏移之间取得平衡。现有研究采用扩散或流生成模型捕捉智能体间的复杂联合行为,但通常依赖多步迭代采样,导致训练与推理效率低下。尽管后续研究通过蒸馏等方法提升采样效率,仍对行为正则化系数敏感。为此,我们提出价值引导多智能体均值流策略(VGM²P),一种简单而高效的基于流的策略学习框架,支持无需调参的条件行为克隆,实现高效动作生成。具体而言,VGM²P利用全局优势值引导智能体协作,将最优策略学习建模为条件行为克隆;同时,为提升多智能体场景下的策略表达能力与推理效率,采用无分类器引导均值流进行策略训练与执行。在具有离散与连续动作空间的任务上,实验表明,即使仅通过条件行为克隆训练,VGM²P也能高效达到与当前最优方法相当的性能。

原文摘要 · Abstract (English)

Offline multi-agent reinforcement learning (MARL) aims to learn the optimal joint policy from pre-collected datasets, requiring a trade-off between maximizing global returns and mitigating distribution shift from offline data. Recent studies use diffusion or flow generative models to capture complex joint policy behaviors among agents; however, they typically rely on multi-step iterative sampling, thereby reducing training and inference efficiency. Although further research improves sampling efficiency through methods like distillation, it remains sensitive to the behavior regularization coefficient. To address the above-mentioned issues, we propose Value Guidance Multi-agent MeanFlow Policy (VGM$^2$P), a simple yet effective flow-based policy learning framework that enables efficient action generation with coefficient-insensitive conditional behavior cloning. Specifically, VGM$^2$P uses global advantage values to guide agent collaboration, treating optimal policy learning as conditional behavior cloning. Additionally, to improve policy expressiveness and inference efficiency in multi-agent scenarios, it leverages classifier-free guidance MeanFlow for both policy training and execution. Experiments on tasks with both discrete and continuous action spaces demonstrate that, even when trained solely via conditional behavior cloning, VGM$^2$P efficiently achieves performance comparable to state-of-the-art methods.

多智能体离线RL流模型策略学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。