arXiv:2510.27123cs.LG2025-10

提出公平性约束的离线强化学习方法,降低不同群体间收益差距。

Group-Sensitive Offline Contextual Bandits

  • 在离线策略优化中加入群体收益差异约束
  • 实测可有效缩小群体间收益差距,且整体性能不降
  • 适合关注算法公平性的资源分配场景

离线上下文老虎机允许从历史数据中学习策略而无需在线交互。然而,仅最大化总体期望回报的离线策略优化可能无意中放大不同群体间的回报差异,导致部分群体获益更多,引发公平性问题,尤其在资源有限时。本文研究了离线上下文老虎机中的群体敏感公平性约束,旨在减少策略学习过程中可能产生的群体间收益差异。针对两类常见公平性要求:将群体间回报差异控制在用户定义的阈值内,或在策略优化中最小化该差异。我们提出一种带约束的离线策略优化框架,将群体间奖励差异约束引入基于离策略梯度的优化过程。为提升训练期间群体间奖励差异的估计精度,采用双重稳健估计器,并提供了策略优化的收敛性保证。在合成和真实数据集上的实验表明,该方法能有效降低奖励差异,同时保持竞争力的总体表现。

原文摘要 · Abstract (English)

Offline contextual bandits allow one to learn policies from historical/offline data without requiring online interaction. However, offline policy optimization that maximizes overall expected rewards can unintentionally amplify the reward disparities across groups. As a result, some groups might benefit more than others from the learned policy, raising concerns about fairness, especially when the resources are limited. In this paper, we study a group-sensitive fairness constraint in offline contextual bandits, reducing group-wise reward disparities that may arise during policy learning. We tackle the following common-parity requirements: the reward disparity is constrained within some user-defined threshold or the reward disparity should be minimized during policy optimization. We propose a constrained offline policy optimization framework by introducing group-wise reward disparity constraints into an off-policy gradient-based optimization procedure. To improve the estimation of the group-wise reward disparity during training, we employ a doubly robust estimator and further provide a convergence guarantee for policy optimization. Empirical results in synthetic and real-world datasets demonstrate that our method effectively reduces reward disparities while maintaining competitive overall performance.

离线强化学习公平性上下文老虎机

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。