arXiv:2603.02333cs.CL2026-03被引 3

揭示扩散语言模型的记忆机制,发现采样精度越高越易复现训练数据。

Characterizing Memorization in Diffusion Language Models: Generalized Extraction and Sampling Effects

  • 提出统一的随机抽取框架,涵盖多种生成模式与掩码策略。
  • 理论证明采样分辨率越高,复现训练数据的概率越高。
  • 实验证明扩散模型比自回归模型更少泄露个人隐私信息。

自回归语言模型(ARMs)被发现会记忆并偶尔完整复现训练数据,引发隐私和版权担忧。扩散语言模型(DLMs)作为新兴替代方案,其记忆行为因生成机制差异仍不明确。本文系统地理论与实证分析了DLMs的记忆特性。提出广义概率抽取框架,统一前缀条件解码与任意掩码模式下的扩散生成。定理4.3建立采样分辨率与记忆之间的单调关系:分辨率越高,精确提取训练数据的概率严格上升,表明自回归解码是扩散生成在最高采样分辨率下的极限情形。跨模型规模与采样策略的大量实验验证了理论预测。在对齐前缀条件评估下,进一步证明DLMs相比ARMs显著降低个人身份信息(PII)的记忆泄露。

原文摘要 · Abstract (English)

Autoregressive language models (ARMs) have been shown to memorize and occasionally reproduce training data verbatim, raising concerns about privacy and copyright liability. Diffusion language models (DLMs) have recently emerged as a competitive alternative, yet their memorization behavior remains largely unexplored due to fundamental differences in generation dynamics. To address this gap, we present a systematic theoretical and empirical characterization of memorization in DLMs. We propose a generalized probabilistic extraction framework that unifies prefix-conditioned decoding and diffusion-based generation under arbitrary masking patterns and stochastic sampling trajectories. Theorem 4.3 establishes a monotonic relationship between sampling resolution and memorization: increasing resolution strictly increases the probability of exact training data extraction, implying that autoregressive decoding corresponds to a limiting case of diffusion-based generation by setting the sampling resolution maximal. Extensive experiments across model scales and sampling strategies validate our theoretical predictions. Under aligned prefix-conditioned evaluations, we further demonstrate that DLMs exhibit substantially lower memorization-based leakage of personally identifiable information (PII) compared to ARMs.

扩散模型记忆机制隐私保护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。