arXiv:2504.21634cs.CYcs.AI2025-04被引 2

用隐私保护的合成数据审计AI公平性,既安全又有效。

Quantitative Auditing of AI Fairness with Differentially Private Synthetic Data

  • 用差分隐私生成模拟真实数据的合成数据
  • 在Adult、COMPAS等数据集上验证公平性指标一致
  • 适合需保护隐私的金融、医疗等领域使用

AI系统公平性审计可识别并量化偏见,但传统方法使用真实数据会引发安全与隐私风险。审计者可能成为敏感信息的保管人,面临网络攻击威胁;即使未发生直接泄露,数据分析也可能间接暴露机密信息。为此,本文提出一种基于差分隐私合成数据的公平性审计框架。通过隐私保护机制生成反映原始数据统计特性的合成数据,兼顾严谨的公平性评估与强隐私保护。在Adult、COMPAS和Diabetes等真实数据集上的实验表明,合成数据与真实数据的公平性指标高度对齐,能有效保留真实数据的公平性特征。结果证明该框架可在保障敏感信息安全的前提下实现有意义的公平性评估,适用于金融、司法、医疗等高敏感领域。

原文摘要 · Abstract (English)

Fairness auditing of AI systems can identify and quantify biases. However, traditional auditing using real-world data raises security and privacy concerns. It exposes auditors to security risks as they become custodians of sensitive information and targets for cyberattacks. Privacy risks arise even without direct breaches, as data analyses can inadvertently expose confidential information. To address these, we propose a framework that leverages differentially private synthetic data to audit the fairness of AI systems. By applying privacy-preserving mechanisms, it generates synthetic data that mirrors the statistical properties of the original dataset while ensuring privacy. This method balances the goal of rigorous fairness auditing and the need for strong privacy protections. Through experiments on real datasets like Adult, COMPAS, and Diabetes, we compare fairness metrics of synthetic and real data. By analyzing the alignment and discrepancies between these metrics, we assess the capacity of synthetic data to preserve the fairness properties of real data. Our results demonstrate the framework's ability to enable meaningful fairness evaluations while safeguarding sensitive information, proving its applicability across critical and sensitive domains.

公平性审计差分隐私合成数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。