研究审计资源有限时企业如何操纵公平性认证,揭示其不可消除的漏洞。
Fairness Auditing: Lower Bounds on Company Manipulation
- 构建公司与审计方的对抗优化模型,模拟资源受限下的公平性审计
- 推导出审计预算、群体不平衡度和容忍度对偏差的理论下限
- 实验证明即使增加审计资源,企业仍可规避检测,适合关注公平性风险的从业者
公平性审计在招聘、信贷等高风险场景中日益成为强制要求。已有研究揭示了黑盒审计的根本不可能性:足够复杂的模型可逃避任何审计策略。本文从反面量化了在有限审计资源下不可避免的审计后操纵程度。将公平性审计建模为计算无界公司与预算受限审计方之间的极小极大优化问题。研究两种审计范式:(i) 固定规模审计集的预算审计方;(ii) 额外要求审计集以 α 误差内估计被认证模型公平性的预算 α-容错审计方。针对两种情形,我们推导出最坏情况下审计后人口均等性偏差的显式下界,其依赖于审计预算、群体不平衡度及公平容忍度。最后,通过线性与神经网络分类器结合简单审计集构造启发式方法,实证展示了这些理论极限。结果表明,增加审计资源虽能缓解操纵空间,但无法彻底消除,凸显有限预算公平认证的根本局限。
原文摘要 · Abstract (English)
Fairness audits are increasingly mandated in high-stakes applications such as hiring, lending, and automated decision-making. Recent work has established fundamental impossibility results for black-box fairness auditing, showing that sufficiently expressive models can evade any auditing strategy. We complement these results by quantifying the extent of unavoidable post-audit manipulation under finite audit resources. We formulate fairness auditing as a min-max optimization between a computationally unbounded company and a budget-constrained auditor. We study two auditing regimes: (i) a budgeted auditor that certifies fairness using a fixed-size audit set, and (ii) a budgeted α-tolerant auditor that additionally requires the audit set to estimate the fairness of the certified model within an α approximation. For both settings, we derive explicit lower bounds on the worst-case post-audit demographic parity deviation as functions of the audit budget, group imbalance, and fairness tolerance. Finally, we empirically illustrate these theoretical limits using simple audit-set construction heuristics with linear and neural network classifiers. Our results demonstrate that increasing audit resources reduces, but does not eliminate, the scope for post-audit manipulation, highlighting fundamental limitations of finite-budget fairness certification.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。