评估生成模型的隐私风险,揭示真实威胁与计算可行性的矛盾
Empirical Privacy Evaluations of Generative and Predictive Machine Learning Models -- A review and challenges for practice
- 通过实证方法检验生成数据的隐私保护效果
- 大样本数据下算法验证有效但对手假设不现实
- 适合关注隐私评估实践的研究者与数据机构
基于机器学习的生成与预测模型,如采用差分隐私等技术训练的合成数据生成器,承诺提供形式化隐私保障,促进敏感数据共享。然而,在部署前必须实证评估生成数据的隐私风险。本文梳理了机器学习生成与预测模型中实证隐私评估的核心概念与假设,并探讨了在包含数百万条记录的场景(如统计机构和医疗提供方)下,生成模型隐私评估面临的实际挑战。研究发现,用于验证训练算法正确性的方法在大数据集上有效,但通常依赖于现实中不合理的攻击者假设。基于此,我们指出评估的计算可行性与威胁模型现实性之间存在关键权衡。最后,提出未来研究的方向与建议。
原文摘要 · Abstract (English)
Synthetic data generators, when trained using privacy-preserving techniques like differential privacy, promise to produce synthetic data with formal privacy guarantees, facilitating the sharing of sensitive data. However, it is crucial to empirically assess the privacy risks associated with the generated synthetic data before deploying generative technologies. This paper outlines the key concepts and assumptions underlying empirical privacy evaluation in machine learning-based generative and predictive models. Then, this paper explores the practical challenges for privacy evaluations of generative models for use cases with millions of training records, such as data from statistical agencies and healthcare providers. Our findings indicate that methods designed to verify the correct operation of the training algorithm are effective for large datasets, but they often assume an adversary that is unrealistic in many scenarios. Based on the findings, we highlight a crucial trade-off between the computational feasibility of the evaluation and the level of realism of the assumed threat model. Finally, we conclude with ideas and suggestions for future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。