arXiv:2605.18809cs.LGcs.AI2026-05被引 2

通过几何投影提升多智能体学习的稳定性与收敛性

Metric-Gradient Projection for Stable Multi-Agent Policy Learning

  • 将联合更新场投影到度量梯度流,消除非势能成分
  • 理论证明其动态具李雅普诺夫函数,平衡差距有显式上界
  • 可作为插件层集成于现有框架,显著提升训练稳定性

一般和多智能体强化学习常受叠加更新场影响,各智能体策略更新会改变其他智能体面临的优化环境。这种耦合可能使集体改进的可积成分与循环交互动态纠缠,导致学习缓慢或不稳定。现有方法如正则化、信用分配和一致性机制通过局部或算法修改来稳定MARL;HPML在此基础上,将联合更新场投影到度量梯度分量。本文提出HPML(Hodge-Projected Multi-agent Learning),将多智能体系统的联合更新场视为$ L^2 $空间中的向量场,并计算其到最近的度量梯度势流的霍奇型投影。HPML沿投影分量进行更新,得到在选定度量和采样测度下最接近的度量梯度场。该投影为变分定义,由泊松型方程表征,并通过基于图和近似神经网络实现,从样本中恢复投影方向。我们证明投影动态具有李雅普诺夫势函数,且平衡差距存在显式加性非势能项。受控实验验证了其几何机制,CTDE基准测试表明,将HPML作为插件投影层使用时,可提升稳定性与归一化回报。

原文摘要 · Abstract (English)

General-sum multi-agent learning is often governed by a stacked update field in which each agent's policy update changes the optimization landscape faced by the others. This coupling can entangle an integrable component of collective improvement with cyclic interaction dynamics, leading to slow or unstable multi-agent learning. Existing approaches, such as regularization, credit assignment, and consensus methods, stabilize MARL through local or algorithmic modifications; HPML complements them by projecting the joint update field onto a metric-gradient component. We introduce \textbf{HPML} (\textbf{H}odge-\textbf{P}rojected \textbf{M}ulti-agent \textbf{L}earning), which views the joint update field of a multi-agent system as an element of an $L^2$ space of vector fields and computes a Hodge-type projection onto the closest metric-gradient potential flow. HPML follows the projected component as the update direction, yielding the closest metric-gradient field under the chosen metric and sampling measure. The projection is defined variationally, characterized by a Poisson-type equation, and implemented through graph-based and amortized neural realizations that recover projected directions from samples. We show that the projected dynamics admit a Lyapunov potential and yield equilibrium-gap bounds with an explicit additive non-potentiality term. Controlled experiments validate the geometric mechanism, and CTDE benchmarks show improved stability and normalized return when HPML is used as a plug-in projection layer in MARL pipelines.

多智能体强化学习稳定性几何投影

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。