arXiv:2511.05518cs.CLcs.AI2025-11EMNLP被引 4

通过制造模型困惑度,高效提取大模型中记忆的训练数据。

Retracing the Past: LLMs Emit Training Data When They Get Lost

  • 设计攻击方法,通过诱导高熵状态暴露记忆内容。
  • 在多个模型上成功提取原文和近似原文数据,无需预先知道训练数据。
  • 揭示对齐模型仍存记忆风险,适合安全与隐私研究者参考。

大语言模型(LLMs)对训练数据的记忆引发严重的隐私与版权问题。现有数据提取方法,尤其是基于启发式的方法,成功率有限且难以揭示记忆泄漏的根本原因。本文提出混淆诱导攻击(CIA),一种系统性提取记忆数据的框架,通过最大化模型不确定性实现。实验表明,模型在产生记忆文本前会出现持续的词元级预测熵上升。CIA通过优化输入片段,刻意诱发这种连续高熵状态。针对对齐模型,进一步提出不匹配监督微调(SFT),同时削弱对齐并诱导目标困惑,提升攻击效果。在多种未对齐与对齐的LLM上测试显示,本方法在无先验知识条件下,能更有效地提取原文及近似原文数据。研究结果凸显各类模型中持久存在的记忆风险,并提供更系统的评估漏洞方法。

原文摘要 · Abstract (English)

The memorization of training data in large language models (LLMs) poses significant privacy and copyright concerns. Existing data extraction methods, particularly heuristic-based divergence attacks, often exhibit limited success and offer limited insight into the fundamental drivers of memorization leakage. This paper introduces Confusion-Inducing Attacks (CIA), a principled framework for extracting memorized data by systematically maximizing model uncertainty. We empirically demonstrate that the emission of memorized text during divergence is preceded by a sustained spike in token-level prediction entropy. CIA leverages this insight by optimizing input snippets to deliberately induce this consecutive high-entropy state. For aligned LLMs, we further propose Mismatched Supervised Fine-tuning (SFT) to simultaneously weaken their alignment and induce targeted confusion, thereby increasing susceptibility to our attacks. Experiments on various unaligned and aligned LLMs demonstrate that our proposed attacks outperform existing baselines in extracting verbatim and near-verbatim training data without requiring prior knowledge of the training data. Our findings highlight persistent memorization risks across various LLMs and offer a more systematic method for assessing these vulnerabilities.

大模型安全记忆泄露隐私攻击

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。