让普通用户也能低成本验证AI是否歧视,只需少量数据和算力。
Audits Under Resource, Data, and Access Constraints: Scaling Laws For Less Discriminatory Alternatives
- 用数学公式推导出公平与性能的最优权衡边界,无需训练大模型。
- 仅需7个小型模型和少量数据,就能预测大型模型的公平性表现。
- 特别适合资源有限的原告方,在信息不透明时证明算法歧视存在。
AI审计在确保人工智能问责与安全中至关重要,尤其涉及反歧视法律中的‘较少歧视性替代方案’(LDA)要求。若某种协议(如模型)不存在显著更少歧视且性能相当的替代方案,则该协议可被辩护。然而,通常由指控方承担举证责任,而其往往缺乏训练高性能低歧视模型所需的资源与数据。此外,开发者常以商业秘密为由限制模型及训练数据的访问。本文提出一种新方法,使指控方在算力、数据、信息和模型访问受限的情况下,仍能判断是否存在满足条件的LDA。聚焦于以人口统计均等性衡量公平性、以二元交叉熵损失衡量性能的场景。我们首次给出损失-公平性帕累托前沿(PF)的闭式上界,并展示如何利用该上界在低资源条件下拟合实际模型的前沿,无需训练任何大型模型。此表达式即为损失-公平性前沿的缩放定律。使用者仅需一小部分训练/测试数据,通过训练最多7个小型模型即可完成上下文特定的前沿拟合。仿真测试表明,该缩放定律在理论假设不完全成立时依然稳健。
原文摘要 · Abstract (English)
AI audits play a critical role in AI accountability and safety. One branch of the law for which AI audits are particularly salient is anti-discrimination law. Several areas of anti-discrimination law implicate the "less discriminatory alternative" (LDA) requirement, in which a protocol (e.g., model) is defensible if no less discriminatory protocol that achieves comparable performance can be found with a reasonable amount of effort. Notably, the burden of proving an LDA exists typically falls on the claimant (the party alleging discrimination). This creates a significant hurdle in AI cases, as the claimant would seemingly need to train a less discriminatory yet high-performing model, a task requiring resources and expertise beyond most litigants. Moreover, developers often shield information about and access to their model and training data as trade secrets, making it difficult to reproduce a similar model from scratch. In this work, we present a procedure enabling claimants to determine if an LDA exists, even when they have limited compute, data, information, and model access. We focus on the setting in which fairness is given by demographic parity and performance by binary cross-entropy loss. As our main result, we provide a novel closed-form upper bound for the loss-fairness Pareto frontier (PF). We show how the claimant can use it to fit a PF in the "low-resource regime," then extrapolate the PF that applies to the (large) model being contested, all without training a single large model. The expression thus serves as a scaling law for loss-fairness PFs. To use this scaling law, the claimant would require a small subsample of the train/test data. Then, the claimant can fit the context-specific PF by training as few as 7 (small) models. We stress test our main result in simulations, finding that our scaling law holds even when the exact conditions of our theory do not.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。