提出新方法识别大模型中存储的个人数据,助力实现删除权请求。
What Should LLMs Forget? Quantifying Personal Data in LLMs for Right-to-Be-Forgotten Requests
- 构建5000+条个人属性语料库,量化模型对个体事实的记忆程度。
- 发现模型记忆强度与个人网络存在度及模型规模正相关。
- 为个人级数据遗忘提供可操作的技术基础,适合隐私合规研究者。
大型语言模型可能记忆并泄露个人信息,引发欧盟GDPR中“被遗忘权”(RTBF)的合规担忧。现有机器遗忘方法假设待遗忘数据已知,但未解决如何识别模型中存储的个体-事实关联。隐私审计技术通常作用于群体层面或少数标识符,难以应对个体级数据查询。本文提出WikiMem,一个包含超过5000条自然语言探针的数据集,覆盖来自Wikidata的243个与人类相关的属性;并设计一种模型无关的度量方法,通过校准负对数似然在改写提示下的表现,对真实值与反事实值进行排序,以量化模型中的人-事实关联。我们在15个不同规模的LLM(参数量410M-70B)上评估了200名个体,发现记忆强度与目标人物的网络存在度及模型规模呈正相关。本工作为在个体层面识别模型中记忆的个人数据提供了基础,支持动态构建遗忘集,适用于机器遗忘和RTBF请求。
原文摘要 · Abstract (English)
Large Language Models (LLMs) can memorize and reveal personal information, raising concerns regarding compliance with the EU's GDPR, particularly the Right to Be Forgotten (RTBF). Existing machine unlearning methods assume the data to forget is already known but do not address how to identify which individual-fact associations are stored in the model. Privacy auditing techniques typically operate at the population level or target a small set of identifiers, limiting applicability to individual-level data inquiries. We introduce WikiMem, a dataset of over 5,000 natural language canaries covering 243 human-related properties from Wikidata, and a model-agnostic metric to quantify human-fact associations in LLMs. Our approach ranks ground-truth values against counterfactuals using calibrated negative log-likelihood across paraphrased prompts. We evaluate 200 individuals across 15 LLMs (410M-70B parameters), showing that memorization correlates with subject web presence and model scale. We provide a foundation for identifying memorized personal data in LLMs at the individual level, enabling the dynamic construction of forget sets for machine unlearning and RTBF requests.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。