arXiv:2601.03791cs.CLcs.AI2026-01ACL被引 3

重新评估大模型泄露个人信息,发现多数是因提示词线索而非真正记忆。

Do LLMs Really Memorize Personally Identifiable Information? Revisiting PII Leakage with a Cue-Controlled Memorization Framework

  • 提出基于提示控制的抗线索记忆评估框架CRM,排除表面线索干扰。
  • 在32种语言中测试,去除线索后信息重建成功率大幅下降。
  • 适合关注大模型隐私安全与评估方法可靠性的研究人员。

大型语言模型(LLMs)被报告存在个人身份信息(PII)泄露问题,通常以成功重建为证据。本文提出对记忆性评估的系统性修正,主张应在低词汇线索条件下评估PII泄露,即目标信息无法通过提示诱导的泛化或模式补全获得。我们形式化了「抗线索记忆」(Cue-Resistant Memorization, CRM)作为必要评估框架,明确控制提示与目标间的重叠线索。利用CRM,在32种语言和多种记忆范式下进行大规模重评估。在回溯型任务(如前后缀完整补全、关联重建)中,发现其有效性主要依赖直接表面形式线索,而非真实记忆;当线索被控制时,重建成功率显著降低。进一步测试无线索生成和成员推断,真阳性率极低。总体表明,先前报告的PII泄露更可能由线索驱动行为导致,而非真实记忆,强调了线索控制对可靠量化大模型隐私相关记忆的重要性。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have been reported to "leak" Personally Identifiable Information (PII), with successful PII reconstruction often interpreted as evidence of memorization. We propose a principled revision of memorization evaluation for LLMs, arguing that PII leakage should be evaluated under low lexical cue conditions, where target PII cannot be reconstructed through prompt-induced generalization or pattern completion. We formalize Cue-Resistant Memorization (CRM) as a cue-controlled evaluation framework and a necessary condition for valid memorization evaluation, explicitly conditioning on prompt-target overlap cues. Using CRM, we conduct a large-scale multilingual re-evaluation of PII leakage across 32 languages and multiple memorization paradigms. Revisiting reconstruction-based settings, including verbatim prefix-suffix completion and associative reconstruction, we find that their apparent effectiveness is driven primarily by direct surface-form cues rather than by true memorization. When such cues are controlled for, reconstruction success diminishes substantially. We further examine cue-free generation and membership inference, both of which exhibit extremely low true positive rates. Overall, our results suggest that previously reported PII leakage is better explained by cue-driven behavior than by genuine memorization, highlighting the importance of cue-controlled evaluation for reliably quantifying privacy-relevant memorization in LLMs.

大模型隐私记忆评估线索控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。