arXiv:2606.18111cs.LGcs.AI2026-06中稿 · the Reinforcement …被引 2

提出多策略公平强化学习方法,让系统自适应不同用户偏好。

Learning Fair Pareto-Optimal Policies in Multi-Objective Reinforcement Learning

  • 用广义吉尼福利函数构建公平性约束,确保策略集覆盖所有偏好
  • 非平稳与随机策略通过历史奖励动态调整,提升分配公平性
  • 适用于需兼顾效率与公平的复杂决策场景,如资源分配

公平性是多目标强化学习(MORL)中决策的重要方面,要求策略在多个可能冲突的目标间实现最优与公平的平衡。现有单策略方法虽能基于固定用户偏好学习公平策略(如使用广义吉尼福利函数,GGF),但无法应对动态或未知的偏好变化。为此,本文形式化了多策略MORL中的公平优化问题:学习一组帕累托最优策略,以覆盖所有可能的用户偏好。关键贡献包括:(1) 证明对于凹形、分段线性的福利函数(如GGF),公平策略始终位于凸覆盖集(CCS)内,即线性加权法的近似帕累托前沿;(2) 证明引入累积奖励历史的非平稳策略和随机策略可动态适应历史不公,提升公平性;(3) 提出三种新算法:将GGF融入多策略多目标Q学习(MOQL)、状态增强型多策略MOQL以学习非平稳策略,及其用于学习随机策略的新扩展。在多个领域评估表明,所提方法能学习出满足不同用户偏好的公平策略集。

原文摘要 · Abstract (English)

Fairness is an important aspect of decision-making in multi-objective reinforcement learning (MORL), where policies must ensure both optimality and equity across multiple, potentially conflicting objectives. While single-policy MORL methods can learn fair policies for fixed user preferences using welfare functions such as the generalized Gini welfare function (GGF), they fail to provide the diverse set of policies necessary for dynamic or unknown user preferences. To address this limitation, we formalize the fair optimization problem in multi-policy MORL, where the goal is to learn a set of Pareto-optimal policies that ensure fairness across all possible user preferences. Our key technical contributions are threefold: (1) We show that for concave, piecewise-linear welfare functions (e.g., GGF), fair policies remain in the convex coverage set (CCS), which is an approximated Pareto front for linear scalarization. (2) We demonstrate that non-stationary policies, augmented with accrued reward histories, and stochastic policies improve fairness by dynamically adapting to historical inequities. (3) We propose three novel algorithms, which include integrating GGF with multi-policy multi-objective Q-Learning (MOQL), state-augmented multi-policy MOQL for learning non-statoinary policies, and its novel extension for learning stochastic policies. We evaluate our algorithms across various domains and compare our methods against the state-of-the-art MORL baselines. The empirical results show that our methods learn a set of fair policies that accommodate different user preferences.

多目标强化学习公平性策略多样性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。