用符号化防护罩实时纠正黑箱决策的偏见,兼顾公平与成本。
Fairness Shields: Safeguarding against Biased Decision Makers
- 部署符号化防护罩持续监控决策流并适时干预
- 在固定周期内保证公平性,且干预成本最低
- 适合需实时公平保障的高风险决策场景
随着基于AI的决策者日益影响人类生活,其对性别、种族等敏感属性的不公平或偏见问题愈发突出。现有大多数偏差防范措施仅提供长期概率公平性保障,可能导致短序列决策中出现具体实例的偏差。本文提出公平屏蔽机制:一个符号化决策者——公平防护罩,持续监控另一部署中的黑箱决策者的行为序列,并通过干预确保满足指定公平准则,同时最小化总干预成本。我们提出了四种计算公平防护罩的算法,其中一种可保证固定时域内的公平性,三种则在固定间隔后保证周期性公平性。在给定未来决策分布及其干预成本的前提下,这些算法求解不同规模的有限时域最优控制问题,具有不同的计算开销和最优性保障。实验表明,该防护罩在多种场景下均能有效确保公平性并保持成本效率。
原文摘要 · Abstract (English)
As AI-based decision-makers increasingly influence human lives, it is a growing concern that their decisions are often unfair or biased with respect to people's sensitive attributes, such as gender and race. Most existing bias prevention measures provide probabilistic fairness guarantees in the long run, and it is possible that the decisions are biased on specific instances of short decision sequences. We introduce fairness shielding, where a symbolic decision-maker -- the fairness shield -- continuously monitors the sequence of decisions of another deployed black-box decision-maker, and makes interventions so that a given fairness criterion is met while the total intervention costs are minimized. We present four different algorithms for computing fairness shields, among which one guarantees fairness over fixed horizons, and three guarantee fairness periodically after fixed intervals. Given a distribution over future decisions and their intervention costs, our algorithms solve different instances of bounded-horizon optimal control problems with different levels of computational costs and optimality guarantees. Our empirical evaluation demonstrates the effectiveness of these shields in ensuring fairness while maintaining cost efficiency across various scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。