发现时间序列大模型中间层存在普遍冗余,可无损删减。
Universal Redundancies in Time Series Foundation Models
- 通过消融实验和对残差流的直接词元归因,分析模型内部机制。
- 所有模型在删除整层后仍保持性能,表明存在广泛冗余。
- 提出基于投影矩阵稳定秩的头消融策略,揭示模型重复模式与季节性偏差来源。
时间序列基础模型(TSFMs)通过大规模预训练实现对未见时间序列的准确预测,无需任务特定微调。我们在多个标准基准上进行大规模评估,发现主流基于Transformer的TSFMs在其中间层存在冗余组件。我们开发了一套用于机制可解释性的工具,包括对特定组件的消融以及对残差流的直接词元归因。研究结果在多种具有不同架构的领先TSFMs及真实世界与合成时间序列数据集上具有一致性。所有模型均对整层消融表现出鲁棒性。此外,我们构建了一个理论框架,将Transformer视为核回归器,提出一种纯内在的头消融策略,基于每头投影矩阵的稳定秩。利用该方法,我们识别出导致广泛观察到的退化现象(如上下文动机的鹦鹉复述、季节性偏差)的具体注意力头。本研究揭示了此类连续时间序列建模架构的普遍特性。
原文摘要 · Abstract (English)
Time Series Foundation Models (TSFMs) leverage extensive pretraining to accurately predict unseen time series during inference, without the need for task-specific fine-tuning. Through large-scale evaluations on standard benchmarks, we find that leading transformer-based TSFMs exhibit redundant components in their intermediate layers. We introduce a set of tools for mechanistic interpretability of TSFMs, including ablations of specific components and direct logit attribution on the residual stream. Our findings are consistent across several leading TSFMs with diverse architectures, and across a diverse set of real-world and synthetic time-series datasets. We discover that all models in our study are robust to ablations of entire layers. Furthermore, we develop a theoretical framework framing transformers as kernel regressors, motivating a purely intrinsic strategy for ablating heads based on the stable rank of the per-head projection matrices. Using this approach, we uncover the specific heads responsible for degenerate phenomena widely observed in TSFMs, such as parroting of motifs from the context and seasonality bias. Our study sheds light on the universal properties of this emerging class of architectures for continuous-time sequence modeling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。