arXiv:2507.14777cs.LG2025-07被引 1

重新审视大模型记忆机制,发现学习不可避免部分记忆,隐私风险或被夸大。

Rethinking Memorization Measures and their Implications in Large Language Models

  • 提出上下文记忆度量,区分记忆与正常语言学习能力。
  • 实验表明优化学习仍会部分记忆训练数据,且不同度量结果不一致。
  • 旧有记忆检测方法可能误判,真正隐私威胁的案例较少。

针对大语言模型(LLM)中的记忆现象及其隐私风险,本文重新审视现有隐私导向的记忆度量方法:基于回忆的记忆度量、反事实记忆度量,以及新提出的上下文记忆度量。将记忆与学习过程中的局部过拟合关联,上下文记忆度量旨在分离记忆行为与模型的上下文学习能力。一个字符串若其因训练导致的回忆超过最优上下文回忆(即无训练时的最佳上下文学习阈值),则被视为上下文记忆。该度量避免了传统回忆度量中‘高回忆即为记忆’的谬误。理论上,上下文记忆度量比反事实记忆度量条件更强。在6个模型家族共18个大模型上,使用多种熵级不同的正式语言进行实验,结果表明:(a) 不同记忆度量对高频字符串的记忆排序存在分歧;(b) 语言的最优学习无法完全避免对训练字符串的部分记忆;(c) 学习性能提升会降低上下文与反事实记忆,但增加基于回忆的记忆度量得分;(d) 以往报告中通过回忆检测出的记忆字符串,大多既无隐私风险,也未被上下文或反事实记忆度量识别。因此,记忆带来的隐私威胁可能被过度强调。

原文摘要 · Abstract (English)

Concerned with privacy threats, memorization in LLMs is often seen as undesirable, specifically for learning. In this paper, we study whether memorization can be avoided when optimally learning a language, and whether the privacy threat posed by memorization is exaggerated or not. To this end, we re-examine existing privacy-focused measures of memorization, namely recollection-based and counterfactual memorization, along with a newly proposed contextual memorization. Relating memorization to local over-fitting during learning, contextual memorization aims to disentangle memorization from the contextual learning ability of LLMs. Informally, a string is contextually memorized if its recollection due to training exceeds the optimal contextual recollection, a learned threshold denoting the best contextual learning without training. Conceptually, contextual recollection avoids the fallacy of recollection-based memorization, where any form of high recollection is a sign of memorization. Theoretically, contextual memorization relates to counterfactual memorization, but imposes stronger conditions. Memorization measures differ in outcomes and information requirements. Experimenting on 18 LLMs from 6 families and multiple formal languages of different entropy, we show that (a) memorization measures disagree on memorization order of varying frequent strings, (b) optimal learning of a language cannot avoid partial memorization of training strings, and (c) improved learning decreases contextual and counterfactual memorization but increases recollection-based memorization. Finally, (d) we revisit existing reports of memorized strings by recollection that neither pose a privacy threat nor are contextually or counterfactually memorized.

大模型记忆机制隐私风险语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。