arXiv:2606.11284cs.MAcs.GT2026-06中稿 · IJCAI

让多智能体在复杂博弈中自动达成高福利协作结果。

Phi-Actor-Critic: Steering General-Sum Games to Pareto-Efficient Correlated Equilibria

论文配图:Phi-Actor-Critic: Steering General-Sum Games to Pareto-Efficient Correlated Equilibria
图 1 · 摘自论文原文
  • 用后悔值最小化引导学习,寻找最优协作解
  • 单次前向传播计算后悔值,大幅降低计算开销
  • 适合资源分配、交通协调等混合动机场景

现实中的多智能体系统(如交通调度、资源分配)常被建模为非零和博弈,个体利益与集体福祉存在冲突。传统深度多智能体强化学习方法难以选择高效均衡,因价值分解受限于单调性假设,策略梯度方法常收敛至稳定但低效的纳什均衡。为此,我们提出Φ-演员-评论家(Φ-AC)框架,利用交换后悔最小化引导学习朝高福利相关均衡(CE)演化。为使反事实后悔估计在深度多智能体强化学习中可行,Φ-AC采用中心化注意力评论家,在一次前向传播中预测向量化的后悔值,避免昂贵的反事实模拟。我们进一步引入基于拉格朗日的均衡选择机制,在保证稳定性的同时优化社会福利。在矩阵博弈、多智能体粒子环境(MPE)及熔炉收获场景中的实验表明,Φ-AC在多种混合动机设置下均能学习到高效且稳定的协调策略,同时保持高集体回报与良好公平性。

原文摘要 · Abstract (English)

Real-world multi-agent systems, from traffic coordination to resource allocation, are often modeled as general-sum games where individual incentives conflict with collective welfare. In these settings, the central challenge is not merely finding an equilibrium, but selecting socially desirable outcomes among many suboptimal Nash equilibria. Standard deep multi-agent reinforcement learning (MARL) methods struggle with this problem, as value-decomposition approaches are constrained by monotonicity assumptions and policy-gradient methods often converge to stable but socially inefficient equilibria. To address this limitation, we propose $Φ$-Actor-Critic ($Φ$-AC), a framework that leverages swap regret minimization to steer learning toward high-welfare correlated equilibria (CE). To make counterfactual regret estimation tractable in deep MARL, $Φ$-AC employs a centralized attention critic that predicts vector-valued regrets in a single forward pass, avoiding computationally expensive counterfactual simulations. We further introduce a Lagrangian-based equilibrium selection mechanism that optimizes social welfare while enforcing stability through regret constraints. Experiments on matrix games, Multi-Agent Particle Environments (MPE), and the Melting Pot Harvest scenario demonstrate that $Φ$-AC learns efficient and stable coordination strategies across diverse mixed-motive settings while maintaining high collective return and competitive fairness.

多智能体博弈论强化学习协同优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。