用广义吉尼系数实现多智能体资源分配的公平性优化
Fair Resource Allocation in Weakly Coupled Markov Decision Processes
- 以广义吉尼函数定义公平性,替代传统总收益目标
- 同质情况下,最优策略具有置换不变性,可简化求解
- 提出基于计数比例的深度强化学习方法,适用于复杂场景
我们研究在弱耦合马尔可夫决策过程(weakly coupled MDP)中如何实现公平的资源分配。多个子MDP原本独立运行,但受资源约束而相互耦合。采用广义吉尼函数定义公平性,而非传统的总和最大化目标。首先给出一个通用但计算昂贵的线性规划解法;随后聚焦于所有子MDP相同的同质情形,首次证明该问题可转化为在置换不变策略类上优化总收益。这一结果使可利用惠特尔指数策略于休息老虎机场景;对于更一般情形,提出一种基于计数比例的深度强化学习方法。通过全面实验验证了理论结果的有效性,证明所提方法能有效实现公平性。
原文摘要 · Abstract (English)
We consider fair resource allocation in sequential decision-making environments modeled as weakly coupled Markov decision processes, where resource constraints couple the action spaces of $N$ sub-Markov decision processes (sub-MDPs) that would otherwise operate independently. We adopt a fairness definition using the generalized Gini function instead of the traditional utilitarian (total-sum) objective. After introducing a general but computationally prohibitive solution scheme based on linear programming, we focus on the homogeneous case where all sub-MDPs are identical. For this case, we show for the first time that the problem reduces to optimizing the utilitarian objective over the class of "permutation invariant" policies. This result is particularly useful as we can exploit Whittle index policies in the restless bandits setting while, for the more general setting, we introduce a count-proportion-based deep reinforcement learning approach. Finally, we validate our theoretical findings with comprehensive experiments, confirming the effectiveness of our proposed method in achieving fairness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。