提出DECAF框架,让多智能体资源分配更公平高效
DECAF: Learning to be Fair in Multi-agent Resource Allocation
- 用双深度Q网络联合优化公平性与效率
- 在多个场景中优于现有公平强化学习方法
- 支持在线灵活调节公平与效率的权衡
众多资源分配问题在中心仲裁者管理资源约束下运行,各智能体评估并传达对资源的偏好。我们将其形式化为分布式评估、集中分配(DECA)问题,并提出在集中式资源分配中学习公平且高效的策略方法。基于双深度Q学习,提出三种方法:(1) 公平性与效用的联合加权优化;(2) 分离优化,分别训练效用与公平性的两个Q估计器;(3) 在线策略扰动,引导已有黑盒效用函数趋向公平解。所提方法在多个资源分配领域表现优于现有公平多智能体强化学习方法,即使使用多种公平性度量也保持优势,并支持灵活的在线公平-效率权衡。
原文摘要 · Abstract (English)
A wide variety of resource allocation problems operate under resource constraints that are managed by a central arbitrator, with agents who evaluate and communicate preferences over these resources. We formulate this broad class of problems as Distributed Evaluation, Centralized Allocation (DECA) problems and propose methods to learn fair and efficient policies in centralized resource allocation. Our methods are applied to learning long-term fairness in a novel and general framework for fairness in multi-agent systems. We show three different methods based on Double Deep Q-Learning: (1) A joint weighted optimization of fairness and utility, (2) a split optimization, learning two separate Q-estimators for utility and fairness, and (3) an online policy perturbation to guide existing black-box utility functions toward fair solutions. Our methods outperform existing fair MARL approaches on multiple resource allocation domains, even when evaluated using diverse fairness functions, and allow for flexible online trade-offs between utility and fairness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。