arXiv:2412.17123cs.LGcs.CY2024-12被引 1

用双模拟度量提升强化学习中的群体公平性

Fairness in Reinforcement Learning with Bisimulation Metrics

  • 基于双模拟度量设计公平性约束的奖励函数与观测动态
  • 在贷款和入学场景中显著减少群体差异
  • 适合关注算法公平性的决策系统研究者

在动态与序列化环境中,确保长期公平性对自动化决策系统至关重要。若智能体仅追求最大化奖励而忽略公平性,可能导致对不同群体或个体的不公对待。本文建立双模拟度量与群体公平性之间的联系,提出一种新方法:利用双模拟度量学习奖励函数与观测动态,使学习者在反映原始问题的同时实现群体公平。我们在包含贷款与大学录取场景的标准公平性基准上进行了实证评估,验证了该方法在缓解序列决策中不公平现象上的有效性。

原文摘要 · Abstract (English)

Ensuring long-term fairness is crucial when developing automated decision making systems, specifically in dynamic and sequential environments. By maximizing their reward without consideration of fairness, AI agents can introduce disparities in their treatment of groups or individuals. In this paper, we establish the connection between bisimulation metrics and group fairness in reinforcement learning. We propose a novel approach that leverages bisimulation metrics to learn reward functions and observation dynamics, ensuring that learners treat groups fairly while reflecting the original problem. We demonstrate the effectiveness of our method in addressing disparities in sequential decision making problems through empirical evaluation on a standard fairness benchmark consisting of lending and college admission scenarios.

强化学习公平性双模拟

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。