用稀疏自编码器揭示时间序列大模型的因果特征层级
Dissecting Chronos: Sparse Autoencoders Reveal Causal Feature Hierarchies in Time Series Foundation Models
- 首次对时间序列大模型应用稀疏自编码器,分析各层特征
- 中层编码器的关键特征导致38.61的预测误差上升,最具因果影响
- 模型依赖突变检测而非周期模式,适合高风险场景可解释性研究
时间序列基础模型(TSFMs)在高风险领域日益广泛应用,但其内部表征仍不透明。本文首次将稀疏自编码器(SAEs)应用于时间序列基础模型,对710M参数的Chronos-T5-Large在六个编码层上的激活进行训练,通过392次单特征消融实验,验证每个被移除特征均导致CRPS指标恶化,确认其因果相关性。分析显示深度依赖特征层次:早期编码层捕捉低频特征,中层集中于关键的突变检测特征,最终层压缩丰富但因果重要性较低的时间概念分类。最关键特征位于中层(单特征最大CRPS增量达38.61),而非语义最丰富的最终层;渐进式消融反而提升预测质量。结果表明机制可解释性可有效迁移至时间序列模型,且Chronos-T5依赖突变检测而非周期模式识别。
原文摘要 · Abstract (English)
Time series foundation models (TSFMs) are increasingly deployed in high-stakes domains, yet their internal representations remain opaque. We present the first application of sparse autoencoders (SAEs) to a TSFM, training TopK SAEs on activations of Chronos-T5-Large (710M parameters) across six layers. Through 392 single-feature ablation experiments, we establish that every ablated feature produces a positive CRPS degradation, confirming causal relevance. Our analysis reveals a depth-dependent hierarchy: early encoder layers encode low-level frequency features, the mid-encoder concentrates causally critical change-detection features, and the final encoder compresses a rich but less causally important taxonomy of temporal concepts. The most critical features reside in the mid-encoder (max single-feature Delta CRPS = 38.61), not in the semantically richest final encoder layer, where progressive ablation paradoxically improves forecast quality. These findings demonstrate that mechanistic interpretability transfers effectively to TSFMs and that Chronos-T5 relies on abrupt-dynamics detection rather than periodic pattern recognition.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。