arXiv:2608.04365cs.LGcs.CR2026-08中稿 · AAAI

用隐私查询技术让模型无法察觉审计重点,防止其作弊操纵公平性结果。

Manipulation-Proof Oblivious Audits against Deceptive Model Providers

  • 通过私有信息检索让审计方能隐蔽查询,模型不知哪些数据被用于审计。
  • 模型若想隐藏不公平行为,必须伪造更多数据,增加被发现概率。
  • 无需修改模型或训练流程,适用于监管审计等真实场景。

审计已成为算法治理的关键工具,为机器学习模型提供外部审查机制。然而,确保评估完整性仍具挑战:在监管场景中,审计常被提前宣告或轻易识别,使模型提供方可操纵过程,无论有意或无意。这一漏洞在公平性评估中尤为严重——提供方可推断敏感属性,并刻意调整不同群体的分配率以满足公平指标。本文提出一种新型审计协议,通过允许审计方以盲态查询模型,显著提升事后对操纵行为的可检测性。该方法利用私有信息检索机制,要求提供方标注大量样本,但无法知晓最终用于审计的具体子集。协议高效,对审计方开销极小,且无需修改被审计模型、训练流程或推理管道。理论分析表明,在此协议下,试图隐藏不公平的提供方必须篡改更大量响应,从而大幅提高操纵的难度与被发现概率。多个典型审计场景的实验验证了该方法的有效性与实用性。

原文摘要 · Abstract (English)

Audits have emerged as a critical instrument for algorithmic governance, providing a mechanism for external scrutiny and governance of machine learning models. However, ensuring the integrity of such assessments remains a challenging issue. For instance in regulatory contexts, audits are typically declared or easily detected, thus enabling model providers to manipulate the process, whether intentionally or inadvertently. This vulnerability is particularly acute in the context of fairness evaluations, in which providers can often infer sensitive attributes and strategically equalize allocation rates between groups to satisfy fairness metrics. In this paper, we introduce a novel audit protocol designed to significantly increase the post-audit detectability of such manipulations by enabling the auditor to query the model in an oblivious manner. Our approach leverages a Private Information Retrieval mechanism to require the provider to label a large set of instances, while preventing it from knowing which subset will ultimately be used for the audit. The protocol is efficient, imposes minimal overhead on the auditor, and requires no modification to the audited model, its training procedure, or its inference pipeline. We provide theoretical guarantees showing that, under this protocol, a provider attempting to hide unfairness must falsify a significantly larger number of responses, thereby increasing both the difficulty and the likelihood of detection of manipulation. Experimental results across representative audit scenarios confirm the effectiveness and practicality of our approach.

算法审计公平性评估隐私计算

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。