arXiv:2412.06966cs.LGcs.AI2024-12NeurIPS被引 14

机器遗忘无法真正删除模型中的数据,对生成式AI监管有重要警示

Machine Unlearning Doesn't Do What You Think: Lessons for Generative AI Policy and Research

  • 提出系统性框架分析遗忘技术与实际目标的偏差
  • 揭示遗忘机制在删除数据或抑制生成物方面存在根本局限
  • 适合政策制定者与研究人员反思生成式AI的可控性设计

机器遗忘被广泛视为解决生成式AI中存在法律或道德问题内容(如隐私、版权、安全等)的方案。例如,用于从模型参数中移除特定个体的个人信息,或清除训练数据中的受版权保护内容。同时,也用于阻止模型生成与特定个体高度相似的内容,或涉及特定概念(如‘蜘蛛侠’)的输出。然而,这些目标——从模型中精准删除信息或抑制特定输出——面临技术和实质性的挑战。本文提供一个框架,帮助机器学习研究者和政策制定者严谨思考这些挑战,识别出遗忘目标与可行实现之间的多重错配。这些错配解释了为何遗忘并非通用解决方案,无法有效约束生成式AI行为以实现更广泛的正面影响。

原文摘要 · Abstract (English)

"Machine unlearning" is a popular proposed solution for mitigating the existence of content in an AI model that is problematic for legal or moral reasons, including privacy, copyright, safety, and more. For example, unlearning is often invoked as a solution for removing the effects of specific information from a generative-AI model's parameters, e.g., a particular individual's personal data or the inclusion of copyrighted content in the model's training data. Unlearning is also proposed as a way to prevent a model from generating targeted types of information in its outputs, e.g., generations that closely resemble a particular individual's data or reflect the concept of "Spiderman." Both of these goals--the targeted removal of information from a model and the targeted suppression of information from a model's outputs--present various technical and substantive challenges. We provide a framework for ML researchers and policymakers to think rigorously about these challenges, identifying several mismatches between the goals of unlearning and feasible implementations. These mismatches explain why unlearning is not a general-purpose solution for circumscribing generative-AI model behavior in service of broader positive impact.

机器遗忘生成式AI政策影响

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。