arXiv:2509.22232cs.LGcs.AI2025-09

提出可透明权衡性能与公平性的强化学习框架,支持多场景公平决策。

Fairness-Aware Reinforcement Learning (FAReL): A Framework for Transparent and Balanced Sequential Decision-Making

  • 构建fMDP模型显式建模个体与群体,定义序列决策中的公平性
  • 在招聘与反欺诈任务中实现更高公平性,性能损失微小
  • 揭示群体与个体公平不互推,适合需双重公平的复杂场景

现实中的序列决策需兼顾公平性。为此,我们提出一种可探索多种性能-公平权衡的框架,使算法能提供透明的决策建议。通过扩展马尔可夫决策过程(fMDP),显式编码个体与群体以捕捉公平性,并在序列决策中形式化公平度量。在两个不同公平需求的场景中评估:招聘任务要求组建强团队且申请人待遇平等;反欺诈任务需精准识别欺诈交易并公平分担客户负担。实验表明,该框架所学策略在多个场景中更公平,仅伴随轻微性能下降。进一步发现,群体公平与个体公平不互相蕴含,凸显框架在需同时满足两类公平时的价值。最后,给出跨场景应用的实践指南。

原文摘要 · Abstract (English)

Equity in real-world sequential decision problems can be enforced using fairness-aware methods. Therefore, we require algorithms that can make suitable and transparent trade-offs between performance and the desired fairness notions. As the desired performance-fairness trade-off is hard to specify a priori, we propose a framework where multiple trade-offs can be explored. Insights provided by the reinforcement learning algorithm regarding the obtainable performance-fairness trade-offs can then guide stakeholders in selecting the most appropriate policy. To capture fairness, we propose an extended Markov decision process, $f$MDP, that explicitly encodes individuals and groups. Given this $f$MDP, we formalise fairness notions in the context of sequential decision problems and formulate a fairness framework that computes fairness measures over time. We evaluate our framework in two scenarios with distinct fairness requirements: job hiring, where strong teams must be composed while treating applicants equally, and fraud detection, where fraudulent transactions must be detected while ensuring the burden on customers is fairly distributed. We show that our framework learns policies that are more fair across multiple scenarios, with only minor loss in performance reward. Moreover, we observe that group and individual fairness notions do not necessarily imply one another, highlighting the benefit of our framework in settings where both fairness types are desired. Finally, we provide guidelines on how to apply this framework across different problem settings.

强化学习公平性序列决策fMDP

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。