利用审计员先验知识,防止平台在公平性审计中作弊
Robust ML Auditing using Prior Knowledge
- 用审计员对任务的先验知识设计防作弊审计方法
- 实验显示可检测出平台隐藏的最大不公平程度
- 适合关注AI监管与公平性审计的研究者
在落实AI监管的诸多技术挑战中,审计操纵是一个关键但未被充分研究的问题:平台可能故意调整对监管者的回应以通过审计,而对其他用户保持不变。本文提出一种基于审计员对任务先验知识的抗操纵审计方法。首先证明,依赖公开先验(如公开数据集)会使审计易受欺骗。随后,形式化了审计员如何利用对真实答案的先验知识来防范操纵。最后,在两个标准数据集上的实验揭示了平台可隐藏的最大不公平程度,仍会被识别为恶意。该工作为更鲁棒的公平性审计开辟了新方向。
原文摘要 · Abstract (English)
Among the many technical challenges to enforcing AI regulations, one crucial yet underexplored problem is the risk of audit manipulation. This manipulation occurs when a platform deliberately alters its answers to a regulator to pass an audit without modifying its answers to other users. In this paper, we introduce a novel approach to manipulation-proof auditing by taking into account the auditor's prior knowledge of the task solved by the platform. We first demonstrate that regulators must not rely on public priors (e.g. a public dataset), as platforms could easily fool the auditor in such cases. We then formally establish the conditions under which an auditor can prevent audit manipulations using prior knowledge about the ground truth. Finally, our experiments with two standard datasets illustrate the maximum level of unfairness a platform can hide before being detected as malicious. Our formalization and generalization of manipulation-proof auditing with a prior opens up new research directions for more robust fairness audits.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。