arXiv:2604.26841cs.LGcs.AI2026-04

语言扩散模型像记忆系统,能找回未见过的数据。

Language Diffusion Models are Associative Memories Capable of Retrieving Unseen Data

论文配图:Language Diffusion Models are Associative Memories Capable of Retrieving Unseen Data
图 1 · 摘自论文原文
  • 将语言扩散模型视为关联记忆,通过条件似然建立记忆吸引子。
  • 训练数据越多,模型从记忆转向泛化,测试样本的吸引子逐渐增强。
  • 用条件熵可检测模型是否在记忆或泛化,适合部署评估。

语言扩散模型何时会记忆训练数据,以及如何定量评估其真正的生成模式?我们发现基于均匀分布的离散扩散模型(UDDMs)本质上表现为具有涌现创造力的关联记忆(AM)。关联记忆的核心是通过在存储数据点周围建立独立吸引盆,可靠恢复为“记忆”。传统模型如霍普菲尔德网络使用显式能量函数保证稳定吸引子;我们拓展此视角,观察到能量并非必需,条件似然最大化也可形成吸引盆。通过评估训练与测试样本的词元恢复能力,我们发现UDDMs存在一个由训练集大小决定的清晰记忆-泛化过渡:随着训练数据增加,训练样本的吸引盆缩小,未见测试样本的吸引盆扩大,最终两者收敛至相同水平。关键在于,仅通过预测词元序列的条件熵即可检测该过渡:记忆状态对应条件熵趋近于零,而泛化状态下多数词元的条件熵保持有限。因此,条件熵为部署模型中记忆-泛化过渡提供了实用探测手段。

原文摘要 · Abstract (English)

When do language diffusion models memorize their training data, and how to quantitatively assess their true generative regime? We address these questions by showing that Uniform-based Discrete Diffusion Models (UDDMs) fundamentally behave as Associative Memories (AMs) $\textit{with emergent creative capabilities}$. The core idea of an AM is to reliably recover stored data points as $\textit{memories}$ by establishing distinct basins of attraction around them. Historically, models like Hopfield networks use an explicit energy function to guarantee these stable attractors. We broaden this perspective by leveraging the observation that energy is not strictly necessary, as basins of attraction can also be formed via conditional likelihood maximization. By evaluating token recovery of $\textit{training}$ and $\textit{test}$ examples, we identify in UDDMs a sharp memorization-to-generalization transition governed by the size of the training dataset: as it increases, basins around training examples shrink and basins around unseen test examples expand, until both later converge to the same level. Crucially, we can detect this transition using only the conditional entropy of predicted token sequences: memorization is characterized by vanishing conditional entropy, while in the generalization regime the conditional entropy of most tokens remains finite. Thus, conditional entropy offers a practical probe for the memorization-to-generalization transition in deployed models.

扩散模型关联记忆条件熵

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。