arXiv:2511.14135cs.LGcs.AI2025-11被引 1

提出自适应公平约束框架,让多智能体协作时自动平衡工作量。

AdaFair-MARL: Enforcing Adaptive Fairness Constraints in Multi-Agent Reinforcement Learning

  • 用可变拉格朗日乘子动态调节公平性约束,避免手动调参。
  • 在医院协调仿真中实现接近100%的公平性满足率,且团队性能不降。
  • 适合需要公平分配任务的协作系统,如医疗调度、分布式计算。

在追求共同目标的异构多智能体系统中,实现公平的工作负载分配仍具挑战性。固定公平惩罚常导致效率低下、训练不稳定及智能体激励冲突。现有公平多智能体强化学习方法通常通过启发式惩罚或标量奖励修改来引入公平性,依赖事后评估,无法保证达到期望的公平水平。为此,本文提出自适应公平多智能体强化学习(AdaFair-MARL)框架,将工作负载公平性显式建模为约束,使智能体在优化团队绩效的同时保持均衡贡献。该框架基于合作马尔可夫博弈,从Jain公平指数(JFI)几何结构推导出公平约束,并证明其可行集可表示为二阶锥形式,从而支持无需人工调参的正规拉格朗日对偶上升更新。在模拟医院协同环境(MARLHospital)中的实验表明,与奖励塑形和固定惩罚方法相比,AdaFair-MARL显著提升工作负载公平性,同时保持团队性能;其公平性约束满足率达0.99-1.00,接近完美。

原文摘要 · Abstract (English)

Fair workload enforcement in heterogeneous multi-agent systems that pursue shared objectives remains challenging. Fixed fairness penalties often introduce inefficiencies, training instability, and conflicting agent incentives. Reward-shaping approaches in fair Multi-Agent Reinforcement Learning (MARL) typically incorporate fairness through heuristic penalties or scalar reward modifications and often rely on post-hoc evaluation. However, these methods do not guarantee that a desired fairness level will be satisfied. To address this limitation, we propose the Adaptive Fairness Multi-Agent Reinforcement Learning (AdaFair-MARL) framework, which formulates workload fairness as an explicit constraint so that agents maintain balanced contributions while optimizing team performance. We present AdaFair-MARL, a constrained cooperative MARL framework whose core algorithmic component is a primal-dual update that enforces workload fairness via adaptive Lagrange multiplier updates. Grounding the framework in a cooperative Markov game, we derive the fairness constraint from Jain's Fairness Index (JFI) geometry and show that the resulting feasible set admits a second-order cone representation, enabling principled Lagrangian dual-ascent updates without manual penalty tuning. Experiments in a simulated hospital coordination environment (MARLHospital) demonstrate the effectiveness of AdaFair-MARL compared to reward-shaping and fixed-penalty fairness methods, improving workload balance while maintaining team performance. We found that AdaFair-MARL achieves nearly perfect constraint satisfaction (0.99-1.00) while significantly improving workload fairness compared to fixed-penalty baselines.

多智能体公平性强化学习约束优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。