首个系统性评估大模型选择性遗忘隐私漏洞的基准测试
Towards Benchmarking Privacy Vulnerabilities in Selective Forgetting with Large Language Models
- 构建首个综合基准,统一评估多种遗忘方法与攻击手段
- 发现不同数据、模型和攻击下隐私泄露差异显著
- 适合关注模型隐私安全的开发者与合规研究人员
人工智能快速发展,模型学习能力不断增强。在关键场景部署中,确保模型隐私与人类价值观对齐至关重要。选择性遗忘(即机器去记忆)作为隐私保护和数据删除的新范式,可让模型有选择地消除已学数据的影响,尤其符合现代数据保护法规并促进价值对齐。然而,该技术在敏感数据场景下引发严重隐私风险。现有遗忘攻击不断涌现,但实验设置各异,导致评估结果过于乐观且不公平,难以客观比较。本文首次提出全面的隐私漏洞评估基准,系统研究多种受害者数据、先进遗忘攻击、遗忘方法及模型架构下的隐私泄露情况,识别影响遗忘隐私泄露的关键因素。基于新发现,我们提供标准化工具,帮助从业者在部署定制化遗忘应用时实现可信的隐私评估。
原文摘要 · Abstract (English)
The rapid advancements in artificial intelligence (AI) have primarily focused on the process of learning from data to acquire knowledgeable learning systems. As these systems are increasingly deployed in critical areas, ensuring their privacy and alignment with human values is paramount. Recently, selective forgetting (also known as machine unlearning) has shown promise for privacy and data removal tasks, and has emerged as a transformative paradigm shift in the field of AI. It refers to the ability of a model to selectively erase the influence of previously seen data, which is especially important for compliance with modern data protection regulations and for aligning models with human values. Despite its promise, selective forgetting raises significant privacy concerns, especially when the data involved come from sensitive domains. While new unlearning-induced privacy attacks are continuously proposed, each is shown to outperform its predecessors using different experimental settings, which can lead to overly optimistic and potentially unfair assessments that may disproportionately favor one particular attack over the others. In this work, we present the first comprehensive benchmark for evaluating privacy vulnerabilities in selective forgetting. We extensively investigate privacy vulnerabilities of machine unlearning techniques and benchmark privacy leakage across a wide range of victim data, state-of-the-art unlearning privacy attacks, unlearning methods, and model architectures. We systematically evaluate and identify critical factors related to unlearning-induced privacy leakage. With our novel insights, we aim to provide a standardized tool for practitioners seeking to deploy customized unlearning applications with faithful privacy assessments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。