系统梳理机器学习中的差分隐私技术演进与应用
Differential Privacy in Machine Learning: A Survey from Symbolic AI to LLMs
- 从符号AI到大模型,梳理差分隐私理论发展脉络
- 分析主流隐私保护训练方法及其实际效果
- 适合关注隐私安全的算法研究者与工程师
机器学习模型不应泄露原本不可访问的特定信息。差分隐私(Differential Privacy, DP)提供了一个正式框架,通过确保任一数据点的加入或移除不会显著改变算法输出,从而限制私人信息的暴露。本文综述了差分隐私的基础定义,并追溯其在关键理论与应用贡献中的演变。接着深入分析差分隐私如何被集成到机器学习模型中,评估现有隐私保护训练方案的有效性。最后,探讨基于差分隐私的机器学习技术在实践中如何进行评估。本工作旨在为安全可信人工智能系统的持续发展提供全面参考。
原文摘要 · Abstract (English)
Machine learning models should not reveal particular information that is not otherwise accessible. Differential privacy provides a formal framework to mitigate privacy risks by ensuring that the inclusion or exclusion of any single data point does not significantly alter the output of an algorithm, thus limiting the exposure of private information. This survey reviews the foundational definitions of differential privacy and traces their evolution through key theoretical and applied contributions. It then provides an in-depth examination of how DP has been integrated into machine learning models, analyzing existing proposals and methods to preserve privacy when training ML models. Finally, it describes how DP-based ML techniques can be evaluated in practice. By offering a comprehensive overview of differential privacy in machine learning, this work aims to contribute to the ongoing development of secure and responsible AI systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。