arXiv:2507.05578cs.LGcs.CL2025-07被引 10

剖析大模型记忆现象的成因、检测与防范方法。

The Landscape of Memorization in LLMs: Mechanisms, Measurement, and Mitigation

  • 分析训练数据重复、微调等导致记忆的关键因素。
  • 评估前缀提取、成员推断等检测技术的有效性。
  • 适合关注隐私安全与模型可控性的研究者参考。

大型语言模型在多项任务中表现出色,但也存在训练数据记忆问题,引发行为可解释性、隐私风险及学习与记忆边界等关切。本文综述近期研究,系统探讨记忆现象的成因、影响因素及检测与缓解方法。重点分析训练数据重复性、训练动态与微调过程对数据记忆的影响。同时评估前缀提取、成员推断和对抗性提示等检测手段的有效性。此外,讨论记忆带来的法律与伦理影响,并提出数据清洗、差分隐私和后训练遗忘等缓解策略。论文揭示了在降低有害记忆与保持模型性能之间权衡的开放挑战,为未来研究指明关键方向。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have demonstrated remarkable capabilities across a wide range of tasks, yet they also exhibit memorization of their training data. This phenomenon raises critical questions about model behavior, privacy risks, and the boundary between learning and memorization. Addressing these concerns, this paper synthesizes recent studies and investigates the landscape of memorization, the factors influencing it, and methods for its detection and mitigation. We explore key drivers, including training data duplication, training dynamics, and fine-tuning procedures that influence data memorization. In addition, we examine methodologies such as prefix-based extraction, membership inference, and adversarial prompting, assessing their effectiveness in detecting and measuring memorized content. Beyond technical analysis, we also explore the broader implications of memorization, including the legal and ethical implications. Finally, we discuss mitigation strategies, including data cleaning, differential privacy, and post-training unlearning, while highlighting open challenges in balancing the need to minimize harmful memorization with model utility. This paper provides a comprehensive overview of the current state of research on LLM memorization across technical, privacy, and performance dimensions, identifying critical directions for future work.

大模型记忆机制隐私保护检测方法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。