arXiv:2511.21514cs.LGcs.AI2025-11被引 4

用机制可解释性解析时间序列Transformer的决策过程

Mechanistic Interpretability for Transformer-based Time Series Classification

  • 将NLP的可解释技术移植到时间序列Transformer中
  • 发现关键注意力头和时间步对分类起决定性作用
  • 适合研究模型内部机理的从业者或可信AI开发者

基于Transformer的模型在时间序列分类等任务中已达到顶尖水平,但其复杂结构使得内部决策机制难以理解。现有可解释方法多聚焦于输入-输出归因,对内部机制仍不透明。本文将激活修补、注意力显著性、稀疏自编码器等机制可解释技术从自然语言处理迁移至专为时间序列设计的Transformer架构中,系统探查各注意力头与时间步的因果角色,揭示模型内部的因果结构。在基准时间序列数据集上的实验构建了信息传播的因果图,明确指出驱动正确分类的关键注意力头与时间位置。此外,展示了稀疏自编码器在发现可解释潜在特征方面的潜力。研究既推动了Transformer可解释性的方法发展,也为时间序列分类中Transformer性能的内在机制提供了新见解。

原文摘要 · Abstract (English)

Transformer-based models have become state-of-the-art tools in various machine learning tasks, including time series classification, yet their complexity makes understanding their internal decision-making challenging. Existing explainability methods often focus on input-output attributions, leaving the internal mechanisms largely opaque. This paper addresses this gap by adapting various Mechanistic Interpretability techniques; activation patching, attention saliency, and sparse autoencoders, from NLP to transformer architectures designed explicitly for time series classification. We systematically probe the internal causal roles of individual attention heads and timesteps, revealing causal structures within these models. Through experimentation on a benchmark time series dataset, we construct causal graphs illustrating how information propagates internally, highlighting key attention heads and temporal positions driving correct classifications. Additionally, we demonstrate the potential of sparse autoencoders for uncovering interpretable latent features. Our findings provide both methodological contributions to transformer interpretability and novel insights into the functional mechanics underlying transformer performance in time series classification tasks.

可解释性时间序列Transformer

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。