arXiv:2512.23943cs.CYcs.LG2025-12

提出可证明搜索到更公平算法的停止策略,避免无休止试错。

Statistical Guarantees in the Search for Less Discriminatory Algorithms

  • 将寻找更公平算法视为最优停止问题,动态判断何时停止搜索
  • 给出高概率上界,证明继续搜索难有更大公平性提升
  • 适用于司法合规场景,尤其适合信贷与住房等高风险决策

美国反歧视法可能要求企业采纳具有相同业务目标但歧视性更低的替代方案(LDA)。近年来研究指出,该原则对就业、信贷和住房等高风险领域的算法决策有直接影响,可能强制企业主动搜索‘更公平的算法’。尽管通过不同随机种子重训练可获得性能相近但歧视性差异显著的模型,企业无法无限重训,关键问题是:何时可证明已尽合理努力?本文将模型多样性下的LDA搜索建模为最优停止问题,提出一种自适应停止算法,能以高概率提供继续训练所能获得的最佳不公平改善上限,使开发者可向法院等机构证明进一步搜索基本无效。我们还展示了在更强分布假设下可得更紧的边界,并在真实信贷与住房数据集上验证了方法有效性。

原文摘要 · Abstract (English)

U.S. discrimination law can impose liability on firms that fail to adopt a less discriminatory alternative (LDA): a decision policy that achieves the same business objectives while reducing disparate impact on legally protected groups. Recent scholarship argues that this doctrine has direct implications for algorithmic decision-making in high-stakes domains such as employment, lending, and housing, potentially obligating firms to search for "less discriminatory algorithms" (Black et al., 2024). Regulators have at times encouraged proactive LDA searches, reinforcing the expectation of a good-faith effort to identify equally performant models with lower disparate impact. Model multiplicity makes such searches plausible: retraining with different random seeds can yield models with comparable predictive performance but materially different disparate impacts. Yet firms cannot retrain indefinitely, raising a central question: when is the search sufficient to demonstrate good faith? We formalize LDA search under multiplicity as an optimal stopping problem in which a developer seeks to produce evidence that further search is unlikely to yield meaningful improvements. Our main contribution is an adaptive stopping algorithm that provides a high-probability upper bound on the best disparate-impact gains attainable through continued retraining, enabling developers to certify (e.g., to a court) that additional search is unlikely to help. We also show how stronger distributional assumptions over the model space can yield tighter bounds, and we validate the approach on real-world credit and housing datasets.

公平算法最优停止合规性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。