arXiv:2507.20708cs.LGmath.OC2025-07被引 2

恶意方可伪造公平假象,揭示审计漏洞

Exposing the Illusion of Fairness: Auditing Vulnerabilities to Distributional Manipulation Attacks

  • 用熵正则与最优传输构建分布投射,生成看似公平的样本
  • 实验证明在标准数据集上可实现合规伪装且统计难辨
  • 适合关注AI监管、公平性审计的从业者参考

AI系统在高风险领域(如欧盟《人工智能法案》定义)的快速部署,凸显了可靠合规审计的重要性。对于二分类器,监管常依赖全局公平指标(如差异影响比)评估歧视风险。在典型审计场景中,被审计方提供数据子集给审计方,监管机构则验证该子集是否代表整体分布。本文研究恶意被审计方能否从非合规原始分布中构造出满足公平约束且外观具代表性的样本,从而制造公平假象。我们将其形式化为受约束的分布投影问题,提出基于熵和最优传输的数学化操纵策略,刻画满足公平要求所需的最小分布偏移。为应对此类攻击,我们通过基于分布距离的统计检验形式化代表性,并系统评估其检测能力。分析揭示了公平操纵在何种条件下可保持统计不可检测,并提供强化监督验证的实践指南。我们在标准表格数据集上验证理论结果,代码已公开于 https://github.com/ValentinLafargue/Inspection。

原文摘要 · Abstract (English)

The rapid deployment of AI systems in high-stakes domains, including those classified as high-risk under the The EU AI Act (Regulation (EU) 2024/1689), has intensified the need for reliable compliance auditing. For binary classifiers, regulatory risk assessment often relies on global fairness metrics such as the Disparate Impact ratio, widely used to evaluate potential discrimination. In typical auditing settings, the auditee provides a subset of its dataset to an auditor, while a supervisory authority may verify whether this subset is representative of the full underlying distribution. In this work, we investigate to what extent a malicious auditee can construct a fairness-compliant yet representative-looking sample from a non-compliant original distribution, thereby creating an illusion of fairness. We formalize this problem as a constrained distributional projection task and introduce mathematically grounded manipulation strategies based on entropic and optimal transport projections. These constructions characterize the minimal distributional shift required to satisfy fairness constraints. To counter such attacks, we formalize representativeness through distributional distance based statistical tests and systematically evaluate their ability to detect manipulated samples. Our analysis highlights the conditions under which fairness manipulation can remain statistically undetected and provides practical guidelines for strengthening supervisory verification. We validate our theoretical findings through experiments on standard tabular datasets for bias detection. Code is publicly available at https://github.com/ValentinLafargue/Inspection.

公平性审计分布操纵合规验证

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。