arXiv:2604.05297cs.AI2026-04

提出多轮价值分解框架,解决多智能体强化学习收敛到次优解的问题。

Breakthrough the Suboptimal Stable Point in Value-Factorization-Based Multi-Agent Reinforcement Learning

  • 通过迭代让劣质动作变为不稳定点,逐步逼近最优策略。
  • 在捕食者-猎物和SMAC任务中性能超越现有方法。
  • 首次揭示次优稳定点是价值分解失败的根源,理论清晰。

价值分解是多智能体强化学习中的主流范式,但其易收敛至次优解的问题长期缺乏理论解释与有效解决。本文提出‘稳定点’这一新理论概念,用于刻画价值分解在一般情况下的收敛特性。分析表明,非最优稳定点是性能下降的主要原因。然而,使最优动作成为唯一稳定点几乎不可行。为此,我们提出多轮价值分解(MRVF)框架:通过衡量相对于先前选择动作的非负收益增量,将低质量动作转化为不稳定点,推动每轮迭代向更优稳定点逼近。在挑战性基准测试(包括捕食者-猎物任务和星战II多智能体挑战,SMAC)上的实验验证了理论分析,并证明了MRVF在性能上优于当前最先进方法。

原文摘要 · Abstract (English)

Value factorization, a popular paradigm in MARL, faces significant theoretical and algorithmic bottlenecks: its tendency to converge to suboptimal solutions remains poorly understood and unsolved. Theoretically, existing analyses fail to explain this due to their primary focus on the optimal case. To bridge this gap, we introduce a novel theoretical concept: the stable point, which characterizes the potential convergence of value factorization in general cases. Through an analysis of stable point distributions in existing methods, we reveal that non-optimal stable points are the primary cause of poor performance. However, algorithmically, making the optimal action the unique stable point is nearly infeasible. In contrast, iteratively filtering suboptimal actions by rendering them unstable emerges as a more practical approach for global optimality. Inspired by this, we propose a novel Multi-Round Value Factorization (MRVF) framework. Specifically, by measuring a non-negative payoff increment relative to the previously selected action, MRVF transforms inferior actions into unstable points, thereby driving each iteration toward a stable point with a superior action. Experiments on challenging benchmarks, including predator-prey tasks and StarCraft II Multi-Agent Challenge (SMAC), validate our analysis of stable points and demonstrate the superiority of MRVF over state-of-the-art methods.

多智能体强化学习价值分解稳定点

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。