arXiv:2411.11315cs.LG2024-11综述被引 94

解析机器遗忘技术如何实现数据删除与隐私保护

A Review on Machine Unlearning

  • 基于数据溯源机制,实现模型对特定数据的可遗忘性
  • 应对GDPR等法规要求,支持用户请求删除训练数据
  • 适合关注数据合规与隐私安全的研究者与开发者

近年来,越来越多的法律法规开始规范用户隐私的使用。例如,《通用数据保护条例》(GDPR)第17条规定的“被遗忘权”要求机器学习应用在用户提出请求时,能够从数据集中移除部分数据并重新训练模型。此外,从安全角度出发,机器学习模型的训练数据(可能包含用户隐私)应得到有效保护,包括适当的清除机制。因此,研究人员提出了多种隐私保护方法来应对此类问题。本文深入回顾了机器学习模型中的安全与隐私挑战。首先,阐述机器学习在日常生活中如何使用用户私有数据,以及GDPR在此问题中的作用。接着,通过描述机器学习模型的安全威胁,介绍机器遗忘的概念及其在保护用户隐私方面的应用。作为论文核心内容,我们系统介绍了当前主流的机器遗忘方法及相关代表性研究成果,并结合数据溯源进行分析。最后,探讨该领域未来的研究挑战。

原文摘要 · Abstract (English)

Recently, an increasing number of laws have governed the useability of users' privacy. For example, Article 17 of the General Data Protection Regulation (GDPR), the right to be forgotten, requires machine learning applications to remove a portion of data from a dataset and retrain it if the user makes such a request. Furthermore, from the security perspective, training data for machine learning models, i.e., data that may contain user privacy, should be effectively protected, including appropriate erasure. Therefore, researchers propose various privacy-preserving methods to deal with such issues as machine unlearning. This paper provides an in-depth review of the security and privacy concerns in machine learning models. First, we present how machine learning can use users' private data in daily life and the role that the GDPR plays in this problem. Then, we introduce the concept of machine unlearning by describing the security threats in machine learning models and how to protect users' privacy from being violated using machine learning platforms. As the core content of the paper, we introduce and analyze current machine unlearning approaches and several representative research results and discuss them in the context of the data lineage. Furthermore, we also discuss the future research challenges in this field.

机器遗忘隐私保护GDPR数据溯源

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。