arXiv:2502.11828cs.LGcs.GT2025-02ICML被引 6

解决高维状态与约束下多群体公平的强化学习问题

Intersectional Fairness in Reinforcement Learning with Large State and Constraint Spaces

  • 基于状态重加权构建多目标优化框架,支持交集群体公平
  • 提出高效算法,可在指数级目标数量下求解
  • 适用于复杂场景,适合关注公平性研究者

传统强化学习旨在最大化期望奖励,但在现实场景中常需同时优化多个目标。例如涉及公平性时,状态可能关联多个(交叉)人口群体,目标是最大化最低收益群体的奖励。本文研究一种多目标优化问题:每个目标由状态相关的单标量奖励函数重加权定义,该设定推广了最小收益群体最大化的任务。我们提出了具有查询效率的算法,即使目标数量呈指数增长,也能在表格型马尔可夫决策过程(MDP)及具备特定结构的大型MDP上求解。最后,通过偏好附加图MDP实验验证理论结果,展示了方法的实际应用价值。

原文摘要 · Abstract (English)

In traditional reinforcement learning (RL), the learner aims to solve a single objective optimization problem: find the policy that maximizes expected reward. However, in many real-world settings, it is important to optimize over multiple objectives simultaneously. For example, when we are interested in fairness, states might have feature annotations corresponding to multiple (intersecting) demographic groups to whom reward accrues, and our goal might be to maximize the reward of the group receiving the minimal reward. In this work, we consider a multi-objective optimization problem in which each objective is defined by a state-based reweighting of a single scalar reward function. This generalizes the problem of maximizing the reward of the minimum reward group. We provide oracle-efficient algorithms to solve these multi-objective RL problems even when the number of objectives is exponentially large-for tabular MDPs, as well as for large MDPs when the group functions have additional structure. Finally, we experimentally validate our theoretical results and demonstrate applications on a preferential attachment graph MDP.

强化学习公平性多目标优化群体公平

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。