arXiv:2601.03087cs.LGcs.CL2026-01ACL被引 6

用主动采样减少40倍查询,高效审计大模型公平性

Audit Me If You Can: Query-Efficient Active Fairness Auditing of Black-Box LLMs

  • 通过构建代理模型版本空间,动态估算公平性指标不确定性
  • 在CivilComments数据集上仅需144次查询即达0.02误差阈值,比传统方法少40倍
  • 适合需要频繁、低成本评估模型公平性的开发者和监管者

大语言模型在不同人群间存在系统性偏差。审计作为黑箱模型应用的问责工具,却面临查询资源消耗大的问题。本文将审计视为目标公平性度量的不确定性估计,提出BAFA——一种高效的主动公平性审计器。BAFA维护一组与已查询得分一致的代理模型,并通过约束经验风险最小化计算公平性指标(如ΔAUC)的置信区间。主动查询策略不断缩小这些区间以降低估计误差。我们在两个标准公平性数据集( extsc{CivilComments} 和 extsc{Bias-in-Bios})上评估,对比分层采样、幂律采样及消融实验。BAFA在严格误差阈值下(如ε=0.02时)相比分层采样最多减少40倍查询(如 extsc{CivilComments}中144对5,956次),随时间表现更优且跨运行方差更低。结果表明,主动采样可显著降低独立公平性审计所需资源,支持持续模型评估。

原文摘要 · Abstract (English)

Large Language Models (LLMs) exhibit systematic biases across demographic groups. Auditing is proposed as an accountability tool for black-box LLM applications, but suffers from resource-intensive query access. We conceptualise auditing as uncertainty estimation over a target fairness metric and introduce BAFA, the Bounded Active Fairness Auditor for query-efficient auditing of black-box LLMs. BAFA maintains a version space of surrogate models consistent with queried scores and computes uncertainty intervals for fairness metrics (e.g., $Δ$ AUC) via constrained empirical risk minimisation. Active query selection narrows these intervals to reduce estimation error. We evaluate BAFA on two standard fairness dataset case studies: \textsc{CivilComments} and \textsc{Bias-in-Bios}, comparing against stratified sampling, power sampling, and ablations. BAFA achieves target error thresholds with up to 40$\times$ fewer queries than stratified sampling (e.g., 144 vs 5,956 queries at $\varepsilon=0.02$ for \textsc{CivilComments}) for tight thresholds, demonstrates substantially better performance over time, and shows lower variance across runs. These results suggest that active sampling can reduce resources needed for independent fairness auditing with LLMs, supporting continuous model evaluations.

公平性审计主动学习大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。