不修改数据就能评估标签隐私,适合大规模系统使用
Observational Auditing of Label Privacy
- 利用数据分布的固有随机性,无需改动原始数据即可审计隐私
- 在Criteo和CIFAR-10上验证了对标签隐私的有效性
- 扩展了传统成员推理,可应用于受保护属性的隐私审计
差分隐私审计对于评估机器学习系统的隐私保障至关重要。然而,现有审计方法在大规模系统中面临挑战,因其需修改训练数据集——例如注入分布外的“诱饵样本”或移除训练样本——这不仅资源消耗大,还带来显著工程开销。本文提出一种新型观测式审计框架,利用数据分布的内在随机性,实现无需更改原始数据的隐私评估。该方法将隐私审计从传统的成员推理扩展至受保护属性(标签为特例),填补了现有技术的关键空白。我们提供了理论基础,并在Criteo和CIFAR-10数据集上进行了实验,证明了其在审计标签隐私保障方面的有效性。该研究为大规模生产环境中的实际隐私审计开辟了新路径。
原文摘要 · Abstract (English)
Differential privacy (DP) auditing is essential for evaluating privacy guarantees in machine learning systems. Existing auditing methods, however, pose a significant challenge for large-scale systems since they require modifying the training dataset -- for instance, by injecting out-of-distribution canaries or removing samples from training. Such interventions on the training data pipeline are resource-intensive and involve considerable engineering overhead. We introduce a novel observational auditing framework that leverages the inherent randomness of data distributions, enabling privacy evaluation without altering the original dataset. Our approach extends privacy auditing beyond traditional membership inference to protected attributes, with labels as a special case, addressing a key gap in existing techniques. We provide theoretical foundations for our method and perform experiments on Criteo and CIFAR-10 datasets that demonstrate its effectiveness in auditing label privacy guarantees. This work opens new avenues for practical privacy auditing in large-scale production environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。