arXiv:2606.30338cs.AIcs.IT2026-06

在只能有限查询模型的情况下,实现高效公平性审计。

Sequential Fairness Auditing with Limited Output Access

论文配图:Sequential Fairness Auditing with Limited Output Access
图 1 · 摘自论文原文
  • 设计可逐步积累证据的序列检验框架,动态决定何时停止审计。
  • 不同公平指标和访问权限下,查询次数差异显著,近阈值场景增益有限。
  • 适用于决策、评分、置信度等多类输出,适合真实部署环境中的独立审计。

外部评估正日益成为AI系统治理的核心。然而,在实际中独立审计者常受限于对已部署模型的访问,只能通过查询交互获取信息。现有公平性评估方法多假设静态数据集与固定样本统计检验,难以适应需在查询约束下逐步收集证据的真实审计场景。本文将公平性审计建模为在有限模型输出访问下的容差感知序列假设检验问题,提出一种序列广义似然比框架,使审计者能在有限审计池中累积证据,并在获得足够合规或违规支持后及时终止。该框架应用于基于决策的统计均等与平等机会审计,当具备更丰富的可观测输出时,进一步扩展至基于分数与对数几率的代理审计。实验表明,公平性度量与模型访问程度显著影响审计效率;丰富输出信息在某些指标与运行条件下可大幅减少查询次数,但在接近阈值情形下收益有限。本工作为现实部署约束下的序列公平性审计提供了实用的统计框架。

原文摘要 · Abstract (English)

External evaluations are becoming increasingly central to the governance of AI systems. In practice, however, independent auditors often have limited access to deployed models and must rely on query-based interactions. Most existing fairness evaluation methods assume static datasets and fixed-sample statistical tests, making them poorly suited to real-world auditing scenarios in which evidence must be collected sequentially under query constraints. In this work, we formulate fairness auditing as a tolerance-aware sequential hypothesis-testing problem under limited model output access. We develop a sequential generalized likelihood-ratio framework that allows auditors to accumulate evidence from a finite audit pool and stop once sufficient support for compliance or violation has been obtained. The framework is instantiated for decision-based Statistical Parity and Equal Opportunity audits, and extended to score- and logit-based proxy audits when richer observables are available. Our results show that both the fairness metric and the level of model access significantly affect audit efficiency, and that the benefits of richer output information are not uniform across auditing settings. In particular, richer outputs can substantially reduce the number of queries required for some fairness metrics and operating regimes, while offering limited gains in near-threshold cases. This work provides a practical statistical framework for sequential fairness auditing under realistic deployment constraints.

公平性审计序列检验查询效率模型访问

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。