arXiv:2410.02005cs.LGstat.ML2024-10被引 4

构建公平性与不确定性评估的基准,提升算法可信度

FairlyUncertain: A Comprehensive Benchmark of Uncertainty in Algorithmic Fairness

  • 提出基于公理的评估框架,确保不确定性估计在不同模型间一致且校准
  • 二分类中简单方法比现有方法更稳定、更准确,回归任务无需显式干预即更公平
  • 适用于关注算法可信赖性与公平性的研究者和实践者

公平预测算法依赖于平等与信任,但现实数据中的固有不确定性挑战了我们做出一致、公平且校准决策的能力。尽管预测误差的公平管理已广泛研究,近年来已有工作开始关注如何公平地处理不可消除的预测不确定性。然而,将不确定性融入公平性的清晰分类体系和明确目标仍不明确。本文提出 FairlyUncertain,一个用于评估公平性中不确定性估计的公理化基准。该基准认为,公平的不确定性估计应在不同学习流程中保持一致,并与观测到的随机性相校准。在十个主流公平性数据集上的大量实验表明:(1) 一种理论合理且简单的二分类不确定性估计方法比以往方法更一致且更校准;(2) 即使改进不确定性估计,拒绝预测虽能降低错误率,却无法缓解群体间的结果不平衡;(3) 在回归任务中引入一致且校准的不确定性估计,可在无需显式公平干预的情况下提升公平性。此外,该基准工具包设计为可扩展、开源,可随领域发展而演进。通过提供标准化框架来评估不确定性与公平性的交互,FairlyUncertain 为更公平、更可信的机器学习实践铺平道路。

原文摘要 · Abstract (English)

Fair predictive algorithms hinge on both equality and trust, yet inherent uncertainty in real-world data challenges our ability to make consistent, fair, and calibrated decisions. While fairly managing predictive error has been extensively explored, some recent work has begun to address the challenge of fairly accounting for irreducible prediction uncertainty. However, a clear taxonomy and well-specified objectives for integrating uncertainty into fairness remains undefined. We address this gap by introducing FairlyUncertain, an axiomatic benchmark for evaluating uncertainty estimates in fairness. Our benchmark posits that fair predictive uncertainty estimates should be consistent across learning pipelines and calibrated to observed randomness. Through extensive experiments on ten popular fairness datasets, our evaluation reveals: (1) A theoretically justified and simple method for estimating uncertainty in binary settings is more consistent and calibrated than prior work; (2) Abstaining from binary predictions, even with improved uncertainty estimates, reduces error but does not alleviate outcome imbalances between demographic groups; (3) Incorporating consistent and calibrated uncertainty estimates in regression tasks improves fairness without any explicit fairness interventions. Additionally, our benchmark package is designed to be extensible and open-source, to grow with the field. By providing a standardized framework for assessing the interplay between uncertainty and fairness, FairlyUncertain paves the way for more equitable and trustworthy machine learning practices.

公平性不确定性基准测试可信赖AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。