arXiv:2606.08810cs.CLcs.LG2026-06

揭秘扩散模型如何从噪声句子嵌入中恢复语言,关键在解码器可读区域。

Continuous Language Diffusion as a Decoder-Interface Problem

论文配图:Continuous Language Diffusion as a Decoder-Interface Problem
图 1 · 摘自论文原文
  • 提出解码器盆地机制,解释为何去噪仅在特定区域可靠。
  • 实测表明解码器能恢复93%-96%的原始词决策,线性读出达97.9%准确率。
  • 揭示公开模型存在界面阶段图谱,适合评估和改进生成系统接口。

高斯噪声污染的句子嵌入虽无直接语义含义,但连续扩散语言模型仍可从中生成流畅文本。我们通过嵌入式语言流(ELF)研究此现象,发现去噪可靠性依赖于轨迹进入解码器可稳定读取词汇的区域——即解码器盆地机制。为此提出诊断协议,评估可去噪性、语义可恢复性、顺序敏感性、解码器兼容性与轨迹可靠性,揭露了传统指标掩盖的缺陷:低均方误差可能丢失语言内容,低困惑度可能反映低熵坍塌,而干净潜在表示可与狭窄解码器盆地共存。解码器边界约束解释了词恢复依赖边际大小与局部解码器敏感性,而非仅依赖潜在误差。对公开的ELF检查点审计揭示了界面相图:早期预测可读性弱,中期出现分歧为竞争区,晚期进入高边际解码器盆地。一旦进入,生成状态上的词实现异常简单:冻结T5的词嵌入查表可恢复93%-96%原解码决策,单个线性读出在32k样本下达到97.9%一致率,残差尾部留下约1.1-1.2的困惑度差距。保守保留门控下,显式诊断监控使边缘规则提前约17%-28%完成去噪。在LangFlow、BitstreamDiffusion及连续潜变量扩散语言模型(Cola-DLM)上验证,界面问题在不同状态对象与解码器下依然成立。因此,连续与潜变量扩散语言模型应作为表示-解码器系统整体评估。

原文摘要 · Abstract (English)

Gaussian-corrupted sentence embeddings have no direct linguistic interpretation, yet continuous diffusion language models can generate fluent text from them. We study this puzzle through Embedded Language Flows (ELF) and identify a decoder-basin mechanism: our evidence suggests that denoising becomes reliable when trajectories reach regions where the native decoder can read stable tokens. We introduce a diagnostic protocol for denoisability, semantic recoverability, order sensitivity, decoder compatibility, and trajectory reliability. It exposes failures hidden by scalar metrics: low mean-squared error can discard linguistic content, low perplexity can reflect low-entropy collapse, and clean latent reconstruction can coexist with a narrow decoder basin. A decoder-margin bound explains why token recovery depends on margin and local decoder sensitivity, not latent error alone. Auditing public ELF checkpoints reveals an interface phase diagram: early predictions are weakly readable, mid-trajectory disagreement marks a competition region, and late predictions enter a high-margin decoder basin. Once inside, token realization is surprisingly simple on generated ELF states: frozen T5 (Text-to-Text Transfer Transformer) token-embedding lookup recovers $93$--$96\%$ of native decoder decisions, and a single linear readout reaches $97.9\%$ agreement at 32k samples, leaving an $\approx1.1$--$1.2$ perplexity gap in a structured residual tail. Under conservative held-out gates, a margin rule exits roughly $17$--$28\%$ earlier in denoising steps under an explicit diagnostic monitor. Boundary checks on LangFlow, BitstreamDiffusion, and the Continuous Latent Diffusion Language Model (Cola-DLM) show that the same interface questions remain meaningful when the state object and decoder change. Continuous and latent diffusion language models should therefore be evaluated as representation-decoder systems.

扩散模型语言生成解码器接口去噪机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。