arXiv:2502.14200cs.AIcs.MA2025-02

提出因果均值场强化学习,让多智能体系统更适应规模变化。

Causal Mean Field Multi-Agent Reinforcement Learning

  • 用结构因果模型捕捉智能体间本质交互关系
  • 在千级智能体环境下训练仍保持稳定性能
  • 适合需要动态扩展智能体数量的场景

多智能体强化学习的可扩展性仍是关键挑战。均值场强化学习(MFRL)通过均值场理论将多智能体问题简化为双智能体问题,缓解了可扩展性问题,但在非平稳环境中缺乏识别关键交互的能力。因果关系蕴含着环境变化下相对稳定的机制。为此,我们提出因果均值场Q学习(CMFQ)算法,在继承MFRL压缩动作-状态空间表示的基础上,显著增强对智能体数量变化的鲁棒性。首先,我们将MFRL决策过程背后的因果关系建模为结构因果模型(SCM);其次,通过干预SCM量化各交互的实质重要性;最后,设计基于因果效应加权的行为信息紧凑表示。在混合合作-竞争博弈与合作博弈中测试,结果表明该方法在大规模智能体训练及超大规模测试中均表现出优异可扩展性。

原文摘要 · Abstract (English)

Scalability remains a challenge in multi-agent reinforcement learning and is currently under active research. A framework named mean-field reinforcement learning (MFRL) could alleviate the scalability problem by employing the Mean Field Theory to turn a many-agent problem into a two-agent problem. However, this framework lacks the ability to identify essential interactions under nonstationary environments. Causality contains relatively invariant mechanisms behind interactions, though environments are nonstationary. Therefore, we propose an algorithm called causal mean-field Q-learning (CMFQ) to address the scalability problem. CMFQ is ever more robust toward the change of the number of agents though inheriting the compressed representation of MFRL's action-state space. Firstly, we model the causality behind the decision-making process of MFRL into a structural causal model (SCM). Then the essential degree of each interaction is quantified via intervening on the SCM. Furthermore, we design the causality-aware compact representation for behavioral information of agents as the weighted sum of all behavioral information according to their causal effects. We test CMFQ in a mixed cooperative-competitive game and a cooperative game. The result shows that our method has excellent scalability performance in both training in environments containing a large number of agents and testing in environments containing much more agents.

多智能体因果推理强化学习可扩展性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。