arXiv:2511.05541cs.CLcs.AI2025-11中稿 · ICLR被引 14

通过时间一致性损失,让自编码器发现更连贯的语义概念。

Temporal Sparse Autoencoders: Leveraging the Sequential Nature of Language for Interpretability

  • 引入时序对比损失,使特征在相邻词上保持一致
  • 无需显式语义信号,仍能提取清晰语义结构
  • 适合研究模型内部表征与可解释性的研究人员

将模型内部表示转化为人类可理解的概念是可解释性的重要目标。尽管近期字典学习方法如稀疏自编码器(SAEs)为发现可解释特征提供了前景,但它们常仅恢复出依赖于具体词元、噪声大或高度局部的概念。我们认为这一局限源于忽视了语言的时间结构——语义内容通常在序列中平滑演变。基于此洞察,我们提出时序稀疏自编码器(T-SAEs),引入一种新颖的对比损失,鼓励高层特征在相邻词元间保持一致激活。这一简单而有效的改进使SAEs能够以自监督方式解耦语义与句法特征。在多个数据集和模型上,T-SAEs恢复出更平滑、更连贯的语义概念,且不牺牲重建质量。令人惊讶的是,即使未使用显式语义信号,其仍展现出清晰的语义结构,为语言模型的无监督可解释性开辟了新路径。

原文摘要 · Abstract (English)

Translating the internal representations and computations of models into concepts that humans can understand is a key goal of interpretability. While recent dictionary learning methods such as Sparse Autoencoders (SAEs) provide a promising route to discover human-interpretable features, they often only recover token-specific, noisy, or highly local concepts. We argue that this limitation stems from neglecting the temporal structure of language, where semantic content typically evolves smoothly over sequences. Building on this insight, we introduce Temporal Sparse Autoencoders (T-SAEs), which incorporate a novel contrastive loss encouraging consistent activations of high-level features over adjacent tokens. This simple yet powerful modification enables SAEs to disentangle semantic from syntactic features in a self-supervised manner. Across multiple datasets and models, T-SAEs recover smoother, more coherent semantic concepts without sacrificing reconstruction quality. Strikingly, they exhibit clear semantic structure despite being trained without explicit semantic signal, offering a new pathway for unsupervised interpretability in language models.

可解释性自编码器语义结构无监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。