现有方法忽略语言的时间动态,新方法捕捉上下文可预测与不可预测信息。
Priors in Time: Missing Inductive Biases for Language Model Interpretability
- 引入时序特征分析,区分上下文可预测与新信息部分
- 在错位句、事件边界等任务上显著优于传统SAE方法
- 适合研究语言模型内部时序表征与动态概念演化
从语言模型激活值中恢复有意义的概念是可解释性研究的核心目标。现有特征提取方法假设概念在时间上相互独立,但这一假设与语言表征的丰富时序结构相悖。通过贝叶斯视角分析发现,稀疏自编码器(SAEs)隐含了概念跨时间独立的先验,暗示平稳性。而语言模型表示具有丰富的时序动态:概念维度系统性增长、上下文依赖相关性明显、显著非平稳。受计算神经科学启发,我们提出一种新的可解释性目标——时序特征分析(Temporal Feature Analysis),其具备时序归纳偏置,能将某一时刻的表示分解为可由上下文推断的部分和残差部分,后者捕捉上下文未解释的新信息。该方法在错位句解析、事件边界识别等任务中表现优异,且能有效分离抽象缓慢变化的信息与快速新信息,而传统SAE在此类任务中存在明显缺陷。结果表明,设计鲁棒可解释工具需匹配数据本身的归纳偏置。
原文摘要 · Abstract (English)
Recovering meaningful concepts from language model activations is a central aim of interpretability. While existing feature extraction methods aim to identify concepts that are independent directions, it is unclear if this assumption can capture the rich temporal structure of language. Specifically, via a Bayesian lens, we demonstrate that Sparse Autoencoders (SAEs) impose priors that assume independence of concepts across time, implying stationarity. Meanwhile, language model representations exhibit rich temporal dynamics, including systematic growth in conceptual dimensionality, context-dependent correlations, and pronounced non-stationarity, in direct conflict with the priors of SAEs. Taking inspiration from computational neuroscience, we introduce a new interpretability objective -- Temporal Feature Analysis -- which possesses a temporal inductive bias to decompose representations at a given time into two parts: a predictable component, which can be inferred from the context, and a residual component, which captures novel information unexplained by the context. Temporal Feature Analyzers correctly parse garden path sentences, identify event boundaries, and more broadly delineate abstract, slow-moving information from novel, fast-moving information, while existing SAEs show significant pitfalls in all the above tasks. Overall, our results underscore the need for inductive biases that match the data in designing robust interpretability tools.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。