发现大模型删数据时,少数群体信息更难彻底删除,隐私泄露风险高出20%以上。
Underestimated Privacy Risks for Minority Populations in Large Language Model Unlearning
- 针对少数群体设计敏感数据探测机制,模拟高风险数据删除场景
- 实验显示少数群体数据隐私泄露率至少高出20%,跨模型、数据集均成立
- 提出新型评估框架,帮助识别现有方法对少数群体的保护盲区
大型语言模型(LLMs)包含敏感的人类生成数据,亟需遗忘(unlearning)方法保障隐私。尽管认证式遗忘提供强隐私保障,但其严格假设不适用于LLMs,导致多数采用启发式方法并依赖经验评估。当前评估通常随机选数据删除后,通过成员推断攻击(MIAs)对比遗忘模型与重新训练模型。然而,为确保每个数据点都获得充分保护,必须考虑某些数据子集面临更高风险的情况。已有研究指出,异常值(尤其是与少数群体相关数据)更易被模型记忆,可能更难遗忘。基于此,我们提出一种补充性的少数群体感知评估框架,以揭示现有框架的盲点。通过精心设计的实验,使用包含个人身份信息(PII)的‘陷阱数据’(canaries)代表少数群体,证明其在多种遗忘方法、MIAs、数据集及模型规模下,隐私泄露至少高出20%。所提框架是实现更公平、全面评估LLM遗忘效果的关键一步。
原文摘要 · Abstract (English)
Large Language Models (LLMs) embed sensitive, human-generated data, prompting the need for unlearning methods. Although certified unlearning offers strong privacy guarantees, its restrictive assumptions make it unsuitable for LLMs, giving rise to various heuristic approaches typically assessed through empirical evaluations. These standard evaluations randomly select data for removal, apply unlearning techniques, and use membership inference attacks (MIAs) to compare unlearned models against models retrained without the removed data. However, to ensure robust privacy protections for every data point, it is essential to account for scenarios in which certain data subsets face elevated risks. Prior research suggests that outliers, particularly including data tied to minority groups, often exhibit higher memorization propensity which indicates they may be more difficult to unlearn. Building on these insights, we introduce a complementary, minority-aware evaluation framework to highlight blind spots in existing frameworks. We substantiate our findings with carefully designed experiments, using canaries with personally identifiable information (PII) to represent these minority subsets and demonstrate that they suffer at least 20% higher privacy leakage across various unlearning methods, MIAs, datasets, and LLM scales. Our proposed minority-aware evaluation framework marks an essential step toward more equitable and comprehensive assessments of LLM unlearning efficacy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。