arXiv:2607.18554cs.MAcs.LG2026-07

提出可扩展的连续多智能体强化学习算法,解决网络化系统中的协同优化问题。

Scalable Policy Optimization for Networked Multi-Agent Reinforcement Learning with Continuous State-Action Spaces

  • 基于局部邻域构建带截断值函数的分布式策略梯度方法
  • 在指数衰减交互条件下,采样复杂度为ε⁻²量级,计算通信仅依赖邻域大小
  • 适用于大规模网络化系统,尤其适合资源受限的分布式智能体场景

我们为具有连续状态与动作空间的网络化马尔可夫决策过程开发了连续分布式耦合策略梯度(CDCPG)算法。每个智能体在其有界邻域内维护本地策略,并通过谱随机特征表示局部转移核,使用最小二乘时差法评估截断动作值函数。分析中提出四项贡献:第一,将截断动作值函数定义为邻域上的条件期望,建立合理的局部贝尔曼理论,避免了简单截断带来的持续核不匹配;第二,揭示归一化随机特征下的时差稳定性维度障碍,证明无条件激励界,将稳定性简化为对称持续激励条件,可通过在线矩阵浓度证书监控;第三,在智能体交互呈指数空间衰减、目标函数光滑且满足激励条件时,CDCPG以˜O(ε⁻²)次共享查询样本,使平均单智能体平稳性度量逼近一个显式表征的近似下界,超出部分低于ε,且误差依赖性匹配非凸一阶最优率;单智能体计算与通信开销仅取决于邻域规模而非全网规模;第四,自适应局部规则选择半径以平衡截断误差与图衰减残差。在一组网络化线性二次基准任务上的实验验证了局部性与特征维数预测。

原文摘要 · Abstract (English)

We develop the Continuous Distributed Coupled Policy Gradient (CDCPG) algorithm for cooperative reinforcement learning in networked Markov decision processes with continuous state and action spaces. Each agent maintains a local actor over a bounded graph neighborhood, and a localized least-squares temporal-difference critic evaluates a truncated action-value function through a spectral random-feature representation of the local transition kernel. The analysis makes four contributions. First, the truncated action-value function is constructed as a conditional expectation over the neighborhood, yielding a well-posed localized Bellman theory that removes the continuation-kernel mismatch of naive truncation arguments. Second, we expose a dimensional obstruction to temporal-difference stability for normalized random features and prove an unconditional excitation bound that reduces stability to a symmetric persistence-of-excitation condition, monitorable through an online matrix-concentration certificate. Third, under exponential spatial decay of agent interactions, the excitation condition, and smoothness of the objective, CDCPG drives an averaged per-agent stationarity measure to within any excess $ε$ of an explicitly characterized approximation floor using $\widetilde{\mathcal{O}}(ε^{-2})$ shared-oracle samples, and the excess dependence matches the smooth nonconvex first-order rate; per-agent computation and communication are governed by the neighborhood size rather than the network size. Fourth, an adaptive-locality rule selects the radius that balances truncation and graph-decay residuals against the target accuracy. Experiments on a networked linear-quadratic benchmark corroborate the locality and feature-dimension predictions.

多智能体强化学习分布式连续控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。