从因果视角出发,构建长期公平决策的评估与平衡框架。
A Causal Lens for Learning Long-term Fair Policies
- 用因果分解方法拆解长期公平性为直接、延迟和虚假影响三部分。
- 发现长期公平性与收益公平性存在内在关联,可统一建模。
- 提出简单有效算法,兼顾短期与长期公平性,适用于动态决策系统。
公平感知学习致力于在存在偏见的数据下设计避免歧视性决策结果的算法。尽管多数研究聚焦于静态场景中的即时偏差,本文强调在动态决策系统中同时考虑长期公平性与瞬时公平性的重要性。在强化学习背景下,我们提出一个通用框架,以不同群体个体可获得的平均期望资格收益差异来衡量长期公平性。通过因果视角,我们将该指标分解为三部分:直接效应、延迟效应以及政策带来的虚假效应。我们分析了这些分量与一种新兴公平概念——收益公平性之间的内在联系,后者旨在控制决策结果的公平性。最后,我们开发了一种简单而有效的平衡多种公平性概念的方法。
原文摘要 · Abstract (English)
Fairness-aware learning studies the development of algorithms that avoid discriminatory decision outcomes despite biased training data. While most studies have concentrated on immediate bias in static contexts, this paper highlights the importance of investigating long-term fairness in dynamic decision-making systems while simultaneously considering instantaneous fairness requirements. In the context of reinforcement learning, we propose a general framework where long-term fairness is measured by the difference in the average expected qualification gain that individuals from different groups could obtain.Then, through a causal lens, we decompose this metric into three components that represent the direct impact, the delayed impact, as well as the spurious effect the policy has on the qualification gain. We analyze the intrinsic connection between these components and an emerging fairness notion called benefit fairness that aims to control the equity of outcomes in decision-making. Finally, we develop a simple yet effective approach for balancing various fairness notions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。