发现微调模型会意外记住输入中的敏感信息,即使这些信息不在训练目标中。
Unintended Memorization of Sensitive Information in Fine-Tuned Language Models
- 设计测试探针,量化模型对仅出现在输入中的敏感信息的意外记忆程度。
- 结果显示,模型规模、语言和任务类型显著影响敏感信息泄露风险。
- 后训练方法在隐私与性能间权衡更稳定,差分隐私在特定场景下效果突出。
在敏感数据上微调大语言模型存在严重的无意记忆和个人身份信息(PII)泄露风险,可能违反隐私法规并危及个体安全。本文系统研究了一种关键且未被充分探索的漏洞:即仅出现在模型输入中而非训练目标中的PII暴露问题。通过合成与真实数据集,我们设计了受控提取探针,量化无意的PII记忆行为,并分析语言、PII频率、任务类型和模型大小等因素对记忆的影响。进一步对比了四种隐私保护方法——差分隐私、机器遗忘、正则化与偏好对齐,评估其在隐私与任务性能间的权衡。结果表明,后训练方法通常提供更一致的隐私-效用平衡;差分隐私在特定设置下显著降低泄露,但可能导致训练不稳定性。这些发现凸显了微调大模型中记忆问题的持续挑战,强调需要更鲁棒、可扩展的隐私保护技术。
原文摘要 · Abstract (English)
Fine-tuning Large Language Models (LLMs) on sensitive datasets carries a substantial risk of unintended memorization and leakage of Personally Identifiable Information (PII), which can violate privacy regulations and compromise individual safety. In this work, we systematically investigate a critical and underexplored vulnerability: the exposure of PII that appears only in model inputs, not in training targets. Using both synthetic and real-world datasets, we design controlled extraction probes to quantify unintended PII memorization and study how factors such as language, PII frequency, task type, and model size influence memorization behavior. We further benchmark four privacy-preserving approaches including differential privacy, machine unlearning, regularization, and preference alignment, evaluating their trade-offs between privacy and task performance. Our results show that post-training methods generally provide more consistent privacy-utility trade-offs, while differential privacy achieves strong reduction in leakage in specific settings, although it can introduce training instability. These findings highlight the persistent challenge of memorization in fine-tuned LLMs and emphasize the need for robust, scalable privacy-preserving techniques.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。