arXiv:2502.15680cs.CLcs.CR2025-02ACL被引 15

模型训练中增删个人数据会引发隐私涟漪效应,导致意外泄露更多敏感信息。

Privacy Ripple Effects from Adding or Removing Personal Information in Language Model Training

  • 后期出现的相似个人信息可触发对早期信息的记忆
  • 添加个人信息会使其他信息记忆度提升7.5倍
  • 删除个人信息反而可能诱发新信息被记住,适合关注隐私安全的研究者

由于个人身份信息(PII)具有敏感性,其所有者有权决定是否将其纳入大语言模型(LLM)训练或要求移除。此外,因数据集更新、重新采集或下游微调阶段引入,PII可能在训练过程中被添加或删除。我们发现,PII的记憶程度是随训练过程动态变化的,且依赖于常见的设计选择。我们识别出三种新现象:(1) 训练后期出现的外观相似的PII会引发对早期序列的记忆,称为辅助记忆,在我们的实验中影响可达1/3;(2) 添加PII会显著提升其他PII的记憶度,最高达约7.5倍;(3) 移除PII反而可能导致其他PII被记憶。模型开发者在训练时应考虑这些一阶与二阶隐私风险,以避免新PII被意外重现。

原文摘要 · Abstract (English)

Due to the sensitive nature of personally identifiable information (PII), its owners may have the authority to control its inclusion or request its removal from large-language model (LLM) training. Beyond this, PII may be added or removed from training datasets due to evolving dataset curation techniques, because they were newly scraped for retraining, or because they were included in a new downstream fine-tuning stage. We find that the amount and ease of PII memorization is a dynamic property of a model that evolves throughout training pipelines and depends on commonly altered design choices. We characterize three such novel phenomena: (1) similar-appearing PII seen later in training can elicit memorization of earlier-seen sequences in what we call assisted memorization, and this is a significant factor (in our settings, up to 1/3); (2) adding PII can increase memorization of other PII significantly (in our settings, as much as $\approx\!7.5\times$); and (3) removing PII can lead to other PII being memorized. Model creators should consider these first- and second-order privacy risks when training models to avoid the risk of new PII regurgitation.

隐私保护模型记忆数据安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。