arXiv:2601.01786cs.LGcs.CR2026-01

针对敏感信息删除难题,提出按风险等级精准遗忘的新型方法。

UnPII: Unlearning Personally Identifiable Information with Quantifiable Exposure Risk

  • 根据个人信息属性的风险等级分级遗忘,而非一刀切
  • 实测遗忘后模型准确率提升11.8%,通用性提高12.4%
  • 可适配企业隐私政策,适合金融医疗等高合规场景

大型语言模型在金融、医疗、政府等关键领域应用日益广泛,引发对训练中敏感个人身份信息(PII)处理的隐私担忧。欧盟《通用数据保护条例》(GDPR)要求响应请求删除PII,亟需可靠且低成本的数据移除方案。机器遗忘成为有前景的定向删数据方向,但现有技术多采用统一遗忘策略,未考虑不同PII属性带来的差异化隐私与业务风险。本文提出首个以PII为中心的遗忘框架UnPII,基于多维风险指标(识別度、敏感度、可用性、关联性、持久性、暴露度、合规性)构建PII风险指数(PRI),实现对各类信息暴露风险的精细评估,并可适配组织隐私策略。我们系统构建了一个包含1,700个实例的合成PII数据集,模拟真实泄露场景。UnPII可无缝集成于梯度上升、负偏好优化、直接偏好优化等主流遗忘算法,不改变其核心原理。实验表明,相比基准方法,UnPII在遗忘过程中平均仅增加27.5%微调开销,同时实现准确率最高提升11.8%、实用性提升6.3%、泛化能力提升12.4%。

原文摘要 · Abstract (English)

The ever-increasing adoption of Large Language Models in critical sectors like finance, healthcare, and government raises privacy concerns regarding the handling of sensitive Personally Identifiable Information (PII) during training. In response, regulations such as European Union's General Data Protection Regulation (GDPR) mandate the deletion of PII upon requests, underscoring the need for reliable and cost-effective data removal solutions. Machine unlearning has emerged as a promising direction for selectively forgetting data points. However, existing unlearning techniques typically apply a uniform forgetting strategy that neither accounts for the varying privacy risks posed by different PII attributes nor reflects associated business risks. In this work, we propose UnPII, the first PII-centric unlearning approach that prioritizes forgetting based on the risk of individual or combined PII attributes. To this end, we introduce the PII risk index (PRI), a composite metric that incorporates multiple dimensions of risk factors: identifiability, sensitivity, usability, linkability, permanency, exposability, and compliancy. The PRI enables a nuanced evaluation of privacy risks associated with PII exposures and can be tailored to align with organizational privacy policies. To support realistic assessment, we systematically construct a synthetic PII dataset (e.g., 1,700 PII instances) that simulates realistic exposure scenarios. UnPII seamlessly integrates with established unlearning algorithms, such as Gradient Ascent, Negative Preference Optimization, and Direct Preference Optimization, without modifying their underlying principles. Our experimental results demonstrate that UnPII achieves the improvements of accuracy up to 11.8%, utility up to 6.3%, and generalizability up to 12.4%, respectively, while incurring a modest fine-tuning overhead of 27.5% on average during unlearning.

隐私保护机器遗忘数据合规风险评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。