arXiv:2504.20927eess.SYcs.LG2025-04中稿 · Learning for Dynam…被引 1

通过精确分解耦合信息,提升多智能体强化学习的样本与计算效率。

Exploiting inter-agent coupling information for efficient reinforcement learning of cooperative LQR

  • 基于智能体间耦合信息,精确分解局部Q函数,避免近似误差。
  • 理论证明其最坏情况样本复杂度等同于集中式方法。
  • 适合需要高效协同控制的机器人、电网调度等场景。

近年来,开发可扩展且高效的多智能体协同控制强化学习算法备受关注。现有研究基于智能体间经验信息结构对局部Q函数进行不精确分解。本文利用智能体间耦合信息,提出一种系统性方法,实现局部Q函数的精确分解。基于该分解,设计了一种近似最小二乘策略迭代算法,并提出两种架构用于学习各智能体的局部Q函数。我们证明了该分解在最坏情况下的样本复杂度与集中式情形相等,并推导出实现更优样本效率的必要充分图结构条件。数值实验表明,该方法在样本效率和计算效率上均有显著提升。

原文摘要 · Abstract (English)

Developing scalable and efficient reinforcement learning algorithms for cooperative multi-agent control has received significant attention over the past years. Existing literature has proposed inexact decompositions of local Q-functions based on empirical information structures between the agents. In this paper, we exploit inter-agent coupling information and propose a systematic approach to exactly decompose the local Q-function of each agent. We develop an approximate least square policy iteration algorithm based on the proposed decomposition and identify two architectures to learn the local Q-function for each agent. We establish that the worst-case sample complexity of the decomposition is equal to the centralized case and derive necessary and sufficient graphical conditions on the inter-agent couplings to achieve better sample efficiency. We demonstrate the improved sample efficiency and computational efficiency on numerical examples.

多智能体强化学习控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。