arXiv:2602.12520cs.LGcs.MA2026-02

用联合状态动作嵌入提升多智能体模型强化学习的规划能力

Multi-Agent Model-Based Reinforcement Learning with Joint State-Action Learned Embeddings

  • 通过变分自编码器学习联合状态动作嵌入,统一表示与想象推演
  • 在星际争霸微操等任务中,相比基线方法提升长期规划性能
  • 适合需要高效数据利用的复杂多智能体协同场景

在部分可观测且高度动态的环境中协调多个智能体,既需信息丰富的表征,又需数据高效的训练。为此,我们提出一种新型基于模型的多智能体强化学习框架,将联合状态-动作表示学习与想象力回放相结合。设计一个基于变分自编码器的世界模型,并引入状态-动作学习嵌入(SALE)。SALE被注入到预测未来轨迹的想象模块,以及通过混合网络整合个体行动价值以估计联合行动值函数的联合智能体网络中。通过将想象轨迹与基于SALE的行动价值相耦合,智能体获得了对其决策如何影响集体结果的更丰富理解,从而在有限真实环境交互下实现更优的长期规划与优化。在星际争霸微管理、多智能体MuJoCo及基于层级的觅食挑战等成熟多智能体基准上的实证研究显示,该方法持续优于基线算法,并验证了联合状态-动作学习嵌入在多智能体基于模型范式中的有效性。

原文摘要 · Abstract (English)

Learning to coordinate many agents in partially observable and highly dynamic environments requires both informative representations and data-efficient training. To address this challenge, we present a novel model-based multi-agent reinforcement learning framework that unifies joint state-action representation learning with imaginative roll-outs. We design a world model trained with variational auto-encoders and augment the model using the state-action learned embedding (SALE). SALE is injected into both the imagination module that forecasts plausible future roll-outs and the joint agent network whose individual action values are combined through a mixing network to estimate the joint action-value function. By coupling imagined trajectories with SALE-based action values, the agents acquire a richer understanding of how their choices influence collective outcomes, leading to improved long-term planning and optimization under limited real-environment interactions. Empirical studies on well-established multi-agent benchmarks, including StarCraft II Micro-Management, Multi-Agent MuJoCo, and Level-Based Foraging challenges, demonstrate consistent gains of our method over baseline algorithms and highlight the effectiveness of joint state-action learned embeddings within a multi-agent model-based paradigm.

多智能体强化学习模型学习嵌入表示

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。