arXiv:2502.17591cs.CL2025-02ICLR被引 11

让大模型主动遗忘敏感信息,保护隐私且不影响使用效果。

Proactive Privacy Amnesia for Large Language Models: Safeguarding PII with Negligible Impact on Model Utility

  • 借鉴认知科学中的遗忘机制,主动识别并删除与隐私相关记忆。
  • 电话号码泄露风险降低100%,地址泄露风险降低9.8%至87.6%。
  • 适合关注大模型隐私安全的研究者与应用开发者。

随着大语言模型(LLMs)的兴起,其在恶意攻击下泄露个人身份信息(PII)的风险日益受到关注。尽管已有研究尝试保护LLM中的PII,但现有方法难以在隐私保护与模型性能之间取得平衡。受认知科学中遗忘现象的启发,本文提出一种新方法——主动隐私遗忘(Proactive Privacy Amnesia, PPA),通过主动识别并清除序列中与PII最相关的记忆,再植入替代记忆以维持模型功能。我们在多个模型上评估了该方法对常见PII(如电话号码、物理地址)的防护能力,针对主流的PII定向攻击进行了测试。结果表明,与现有防御技术相比,本方法可完全消除电话号码泄露风险(100%),显著降低物理地址泄露风险(9.8%–87.6%),同时保持与原始模型相当的性能表现。

原文摘要 · Abstract (English)

With the rise of large language models (LLMs), increasing research has recognized their risk of leaking personally identifiable information (PII) under malicious attacks. Although efforts have been made to protect PII in LLMs, existing methods struggle to balance privacy protection with maintaining model utility. In this paper, inspired by studies of amnesia in cognitive science, we propose a novel approach, Proactive Privacy Amnesia (PPA), to safeguard PII in LLMs while preserving their utility. This mechanism works by actively identifying and forgetting key memories most closely associated with PII in sequences, followed by a memory implanting using suitable substitute memories to maintain the LLM's functionality. We conduct evaluations across multiple models to protect common PII, such as phone numbers and physical addresses, against prevalent PII-targeted attacks, demonstrating the superiority of our method compared with other existing defensive techniques. The results show that our PPA method completely eliminates the risk of phone number exposure by 100% and significantly reduces the risk of physical address exposure by 9.8% - 87.6%, all while maintaining comparable model utility performance.

隐私保护大模型遗忘机制PII

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。