提出新框架统一刻画大规模多智能体强化学习的局部性,提升理论精度与适用范围。
A Unified Framework for Locality in Scalable MARL
- 分解状态与动作敏感度,构建更精细的局部性分析矩阵
- 证明谱半径小于1可保证值函数衰减,比传统方法更宽松有效
- 适用于软最大化策略,温度参数直接调控系统局部性
在平均奖励设定下,可扩展的网络化多智能体强化学习依赖于系统的值局部性:远距离智能体间的扰动对长期价值影响微弱。传统方法通过单一矩阵 $C^π$ 的 Dobrushin 行和界来验证局部性,但常因取所有联合动作上界而过于保守。本文将 $C^π$ 分解为环境敏感度 $E^{ m s}$ 与策略敏感度 $E^{ m a}Π(π)$,其中 $Π(π)$ 衡量策略对状态变化的响应程度。定义 $H^π = E^{ m s} + E^{ m a}Π(π)$,其谱半径 $ρ(H^π)<1$ 可控制平均奖励泊松解的衰减,该条件严格弱于传统行和界 $\|H^π\|_\infty<1$,且在策略不常选择最坏动作时更具优势。对于温度 $τ$ 的软最大化策略,有 $Π(π) ≤ L/(2τ)$,表明温度越高,系统越局部。基于此衰减结果,本文给出一种块坐标 KL 近端策略改进模板的确定性代理保证,其截断偏差随消息传递半径 $κ$ 指数级衰减。
原文摘要 · Abstract (English)
Scalable methods for networked multi-agent reinforcement learning let each agent plan using only a small neighborhood of the agent graph. This works only when the system is value-local, meaning a perturbation at one agent affects the long-run value at another agent weakly when the two are far apart. In the average-reward setting, the standard way to certify locality is the Dobrushin row-sum bound on a single matrix $C^π$ that captures how each agent's next state depends on each other agent's current state. To make this matrix easy to work with, prior work bounds it by a supremum over joint actions. The resulting bound is independent of the policy, but it is loose whenever the policy never picks the worst-case action. We split $C^π$ into pieces that separately track environment sensitivity and policy sensitivity, $C^π\preceq E^{\mathrm s}+E^{\mathrm a}Π(π)$, where $E^{\mathrm s}$ measures how the next state moves with the current state, $E^{\mathrm a}$ measures how it moves with the current action, and $Π(π)$ measures how reactive the policy is to changes in state. The spectral radius of $H^π:= E^{\mathrm s}+E^{\mathrm a}Π(π)$ then controls the decay of the average-reward Poisson solution, and the spectral certificate $ρ(H^π)<1$ is strictly weaker than the row-sum condition $\|H^π\|_\infty<1$ on the same matrix and applies in regimes where policy-independent action-supremum bounds used in prior Dobrushin-style work cannot. For temperature-$τ$ softmax policies we get $Π(π)\le L/(2τ)$, so the softmax temperature directly controls locality. We use this decay result to give a deterministic oracle guarantee for a block-coordinate KL-proximal policy-improvement template whose truncation bias decays exponentially in the message-passing radius $κ$.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。