arXiv:2607.01774cs.AIcs.CL2026-07

发现扩散语言模型内部隐含时间进度信号,可调控模型信心与熵。

Subliminal Clocks: Latent Time Modelling in Diffusion Language Models

论文配图:Subliminal Clocks: Latent Time Modelling in Diffusion Language Models
图 1 · 摘自论文原文
  • 在残差流中探测到与扩散步数相关的潜在表示
  • 通过低维子空间调节可系统改变模型置信度和熵
  • 该信号具有结构化特征,利于解释模型内部机制

扩散语言模型(DLMs)作为自回归模型的新兴替代方案受到关注。与标准扩散方法不同,DLMs 不显式依赖时间步,引发一个关键问题:模型是否内含去噪进度信息?本文证明,DLMs 确实在其残差流中编码了与扩散时间步相关的潜在表示。我们通过跨层探针可靠地提取该信号,表明去噪进度可从内部激活中解码。进一步实验显示,沿推断时间步对应的低维子空间进行引导,可系统调节模型对去噪进度的认知,从而导致置信度与熵的可预测变化。最后,我们分析了该表示的几何特性,发现其在激活空间中具有结构化且可解释的性质,揭示了此类信号如何被模型处理。

原文摘要 · Abstract (English)

Diffusion Language Models (DLMs) have recently emerged as a promising alternative to autoregressive models. Unlike standard diffusion-based approaches, DLMs are not explicitly conditioned on a timestep, raising a natural question: do these models internally represent denoising progress, and how is such information used downstream? In this work, we show that DLMs do in fact encode a latent representation related to the diffusion timestep within their residual streams. We find that this signal can be reliably extracted using probes across layers, indicating that denoising progress is decodable from internal activations. We further demonstrate that steering the model along a low-dimensional subspace associated with the inferred timestep allows us to systematically modulate its notion of denoising progress, leading to predictable changes in model confidence and entropy. Finally, we analyse the geometry of the identified representation, showing that it exhibits structured and interpretable properties in activation space, and shedding light on how such a signal is processed by these models.

扩散模型语言模型隐变量可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。