arXiv:2607.17117cs.LGcs.CL2026-07

让语言模型的特征具备时间延续性,更好捕捉语义演化。

Persistent Sparse Autoencoders: Learning Feature Timescales in Language Models

论文配图:Persistent Sparse Autoencoders: Learning Feature Timescales in Language Models
图 1 · 摘自论文原文
  • 为每个特征引入持久系数,学习其在序列中持续的时间长度。
  • 慢速特征能保留主题信息,快速特征则检测局部语义变化。
  • 适合用于长文本监控与可解释性分析,如对抗提示注入检测。

稀疏自编码器(SAEs)将语言模型激活值分解为稀疏特征,但标准SAE对每个词元独立编码,无法体现跨序列的持续信息。我们提出持久稀疏自编码器(Persistent SAEs),通过为每个特征学习持久系数,使模型能够识别哪些特征应持续存在及其持续时长。实验表明,该方法在保持竞争性重建质量的同时,学习到一系列特征时间尺度:快速特征表现为局部可解释的检测器,而慢速特征则在持续状态中集中主题级信息。此外,在提示注入监测案例研究中,慢速特征能长期保持检测信号并维持因果有效性。结果表明,持久SAEs为通过持久语义表示解释和监控语言模型开辟了新路径。

原文摘要 · Abstract (English)

Sparse autoencoders (SAEs) decompose language model activations into sparse features, but standard SAEs encode each token independently and do not expose information that persists across a sequence. We introduce Persistent Sparse Autoencoders (Persistent SAEs), which extend standard SAEs by learning a persistence coefficient for each feature, allowing the model to learn which features should persist and for how long. Our experiments show that they retain competitive reconstruction quality while learning a spectrum of feature timescales: fast features behave as locally interpretable detectors, whereas slow features concentrate topic-level information in a persistent state. Moreover, as shown in a prompt-injection monitoring case study, slow features preserve detection signals and remain causally effective over long contexts. These results suggest that Persistent SAEs open up new opportunities for interpreting and monitoring language models through persistent semantic representations.

稀疏编码语义持久性可解释性语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。