arXiv:2604.23235cs.CL2026-04被引 1

通过扩散模型的去噪过程,揭示语言信息在生成中何时浮现。

Measuring Temporal Linguistic Emergence in Diffusion Language Models

论文配图:Measuring Temporal Linguistic Emergence in Diffusion Language Models
图 1 · 摘自论文原文
  • 利用去噪轨迹分析信息出现时机,设计四类时序测量方法。
  • 语义类别比词性更早稳定,粗粒度标签比词汇身份更易恢复。
  • 中段去噪阶段最敏感,适合研究模型内部状态演化与干预效果。

扩散语言模型具有显式的去噪轨迹,使我们能够探究不同信息类型在生成过程中何时变得可测量。本文对LLaDA-8B-Base在掩码WikiText-103文本上进行三次独立的32步运行,每轮包含1,000个探针训练序列和200个保留评估序列。从保存的轨迹中提取四类时序测量:词元承诺、词性(POS)、粗粒度语义类别和词元身份的线性可恢复性、置信度与熵动态变化,以及中段重掩码下的敏感性。跨种子实验显示一致顺序:内容类标签比功能密集类更早稳定;在探针设置下,POS与粗粒度语义标签的线性可恢复性显著高于精确词元身份;最终错误的词元在整个轨迹中不确定性更高,尽管后期置信度校准性下降;扰动敏感性在轨迹中段达到峰值。直接/旁路分解表明该峰值几乎完全局限于扰动位置本身。在该设置下,去噪时间成为有效的分析轴:粗粒度标签较早且更稳健地被恢复,轨迹级不确定性追踪最终正确性,中段状态最具干预敏感性。

原文摘要 · Abstract (English)

Diffusion language models expose an explicit denoising trajectory, making it possible to ask when different kinds of information become measurable during generation. We study three independent 32-step runs of LLaDA-8B-Base on masked WikiText-103 text, each with 1{,}000 probe-training sequences and 200 held-out evaluation sequences. From saved trajectories, we derive four temporal measurements: token commitment; linear recoverability of part-of-speech (POS), coarse semantic category, and token identity; confidence and entropy dynamics; and sensitivity under mid-trajectory re-masking. Across seeds, the same ordering recurs: content categories stabilize earlier than function-heavy categories, POS and coarse semantic labels remain substantially more linearly recoverable than exact lexical identity under our probe setup, uncertainty remains higher for tokens that ultimately resolve incorrectly even though late confidence becomes less calibrated, and perturbation sensitivity peaks in the middle of the trajectory. A direct/collateral decomposition shows that this peak is overwhelmingly local to the perturbed positions themselves. In this LLaDA+WikiText setting, denoising time is therefore a useful analysis axis: under our measurements, coarse labels are recovered earlier and more robustly than lexical identity, trajectory-level uncertainty tracks eventual correctness, and mid-trajectory states are the most intervention-sensitive.

扩散模型语言生成时序分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。