arXiv:2602.05859cs.LGcs.AI2026-02被引 5

首个面向扩散语言模型的可解释性框架,揭示其独特机制。

DLM-Scope: Mechanistic Interpretability of Diffusion Language Models via Sparse Autoencoders

  • 用稀疏自编码器提取扩散模型的可读特征
  • 早期层插入编码器可降低交叉熵损失,优于传统语言模型
  • 适用于扩散时间干预与解码顺序分析,适合研究者探索新方向

稀疏自编码器(SAEs)已成为自回归大语言模型(LLMs)机制可解释性的标准工具,能提取稀疏且人类可读的特征,并干预模型行为。随着扩散语言模型(DLMs)日益成为自回归模型的有力替代,亟需为这类新兴模型开发定制化的可解释性工具。本文提出 DLM-Scope,首个基于 SAE 的 DLM 可解释性框架,证明训练后的 Top-K SAE 能忠实提取可解释特征。值得注意的是,SAE 插入对 DLM 的影响不同于自回归模型:在早期层应用时,可降低交叉熵损失,这一现象在 LLM 中不存在或微弱得多。此外,DLM 中的 SAE 特征支持更有效的扩散时间干预,常优于 LLM 驱动方法。我们还首次探索了 SAE 在 DLM 中的新方向:其可提供解码顺序的有用信号;且在模型后训练阶段特征保持稳定。本工作为 DLM 的机制可解释性奠定基础,并展示将 SAE 应用于 DLM 相关任务的巨大潜力。

原文摘要 · Abstract (English)

Sparse autoencoders (SAEs) have become a standard tool for mechanistic interpretability in autoregressive large language models (LLMs), enabling researchers to extract sparse, human-interpretable features and intervene on model behavior. Recently, as diffusion language models (DLMs) have become an increasingly promising alternative to the autoregressive LLMs, it is essential to develop tailored mechanistic interpretability tools for this emerging class of models. In this work, we present DLM-Scope, the first SAE-based interpretability framework for DLMs, and demonstrate that trained Top-K SAEs can faithfully extract interpretable features. Notably, we find that inserting SAEs affects DLMs differently than autoregressive LLMs: while SAE insertion in LLMs typically incurs a loss penalty, in DLMs it can reduce cross-entropy loss when applied to early layers, a phenomenon absent or markedly weaker in LLMs. Additionally, SAE features in DLMs enable more effective diffusion-time interventions, often outperforming LLM steering. Moreover, we pioneer certain new SAE-based research directions for DLMs: we show that SAEs can provide useful signals for DLM decoding order; and the SAE features are stable during the post-training phase of DLMs. Our work establishes a foundation for mechanistic interpretability in DLMs and shows a great potential of applying SAEs to DLM-related tasks and algorithms.

扩散模型可解释性自编码器语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。