arXiv:2606.13076cs.MAcs.GT2026-06

提出兼顾公平与效率的多智能体强化学习框架,解决合作中的不平等问题。

$α$-fair heterogeneous agent reinforcement learning

  • 用动态加权优势函数实现从效率到公平的平滑过渡
  • 算法在清洁任务中比基线提升12%效率并改善社会收益
  • 适合需要公平合作的复杂多智能体系统研究者

多智能体系统中的协作通常通过最大化整体效率的功利目标优化,但忽视了奖励分配,常导致不平等的“领导者-追随者”关系。尽管基于公平性的方法能促进互利合作,但许多现有算法(包括使用奖励塑形的方法)破坏了马尔可夫博弈的平稳性,或缺乏严格的理论保证。本文提出一种新框架,将α-公平性与异质智能体信任区域学习(HATRL)结合,确保单调改进并收敛至纳什均衡。该方法引入公平优势函数,根据各智能体预期回报动态加权其效用,使全局目标可由纯功利效率渐变为α-公平福利,取决于参数α。我们设计了两种实用算法:α-公平HATRPO和α-公平HAPPO。在序列社会困境任务如CleanUp和CommonHarvest上的实验表明,这些算法在功利视角下表现优于原有HATRL方法,同时实现更高社会收益。

原文摘要 · Abstract (English)

Cooperation in multi-agent systems is typically optimized through utilitarian objectives that maximize overall efficiency but fail to account for reward distribution, often resulting in inequitable "leader-follower" dynamics. While fairness-based approaches encourage pro-social behaviors where every agent benefits from cooperation, many current algorithms - including those utilizing reward shaping - break the stationarity of Markov Games or lack rigorous theoretical guarantees. This creates a critical gap between fair objective methods and theoretically safe learning frameworks. We propose a novel framework that bridges $α$-fairness with Heterogeneous-Agent Trust Region Learning (HATRL), ensuring monotonic improvement and convergence toward Nash Equilibria. Our approach leverages a fair advantage function that dynamically weights agent utilities based on their expected returns, allowing the global objective to transition from purely utilitarian efficiency to $α$-fairness welfare based on the parameter $α$. We introduce two practical algorithms, $α$-fair HATRPO and $α$-fair HAPPO, and demonstrate through experiments in sequential social dilemmas like CleanUp and CommonHarvest that they perform better than HATRL's algorithms from a utilitarian point of view while achieving socially higher outcomes.

多智能体公平性强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。