提出多任务强化学习公平性新方法,确保不同群体间结果更均衡。
Group Fairness in Multi-Task Reinforcement Learning
- 设计约束优化算法,同时在多个任务中强制公平性约束。
- 实验显示公平差距更小,且各群体回报接近,无显著性能损失。
- 适用于需兼顾多任务公平性的现实场景,如智能医疗、推荐系统。
本文针对强化学习(RL)应用中的关键社会议题——在多任务设置下确保不同人口群体间结果的公平性。尽管已有研究关注单任务强化学习的公平性,但许多真实应用场景具有多任务特性,要求策略在所有任务中均保持公平。本文提出一种新的多任务群体公平性形式化定义,并设计了一种约束优化算法,可同时在多个任务中显式施加公平性约束。理论证明该算法在有限时域周期设定下以高概率不违反公平性约束,且具有次线性遗憾。在RiverSwim和MuJoCo环境中的实验表明,相比缺乏显式多任务公平性约束的先前方法,本方法在有限时域与无限时域设定下均能更好地保障多任务公平性。结果显示,所提算法在维持各群体与任务间相当回报水平的同时,显著缩小了公平差距,展现出在真实多任务强化学习场景中缓解公平问题的潜力。
原文摘要 · Abstract (English)
This paper addresses a critical societal consideration in the application of Reinforcement Learning (RL): ensuring equitable outcomes across different demographic groups in multi-task settings. While previous work has explored fairness in single-task RL, many real-world applications are multi-task in nature and require policies to maintain fairness across all tasks. We introduce a novel formulation of multi-task group fairness in RL and propose a constrained optimization algorithm that explicitly enforces fairness constraints across multiple tasks simultaneously. We have shown that our proposed algorithm does not violate fairness constraints with high probability and with sublinear regret in the finite-horizon episodic setting. Through experiments in RiverSwim and MuJoCo environments, we demonstrate that our approach better ensures group fairness across multiple tasks compared to previous methods that lack explicit multi-task fairness constraints in both the finite-horizon setting and the infinite-horizon setting. Our results show that the proposed algorithm achieves smaller fairness gaps while maintaining comparable returns across different demographic groups and tasks, suggesting its potential for addressing fairness concerns in real-world multi-task RL applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。