arXiv:2410.06070cs.LG2024-10被引 5

用概念瓶颈框架让时间序列Transformer模型更可解释。

Interpretability for Time Series Transformers using A Concept Bottleneck Framework

  • 通过中心核对齐约束,引导模型学习预设可解释概念。
  • 模型性能基本不变,但可解释性显著提升。
  • 适合关注模型决策逻辑的研究者与工程师。

机制可解释性致力于逆向解析神经网络内部学习的机制。本文提出一种基于概念瓶颈模型的正向工程框架,聚焦于长期时间序列预测任务。通过修改训练目标,利用中心核对齐(Centered Kernel Alignment)促使模型表示与预定义的可解释概念保持一致,从而引导瓶颈组件学习这些概念,同时允许其他组件学习未定义的概念。该方法应用于Vanilla Transformer、Autoformer和FEDformer,在合成数据及多个基准数据集上进行了深入分析。结果表明,模型性能基本不受影响,而可解释性显著增强。此外,通过激活修补干预实验验证了瓶颈组件的解释有效性。

原文摘要 · Abstract (English)

Mechanistic interpretability focuses on reverse engineering the internal mechanisms learned by neural networks. We extend our focus and propose to mechanistically forward engineer using our framework based on Concept Bottleneck Models. In the context of long-term time series forecasting, we modify the training objective to encourage a model to develop representations which are similar to predefined, interpretable concepts using Centered Kernel Alignment. This steers the bottleneck components to learn the predefined concepts, while allowing other components to learn other, undefined concepts. We apply the framework to the Vanilla Transformer, Autoformer and FEDformer, and present an in-depth analysis on synthetic data and on a variety of benchmark datasets. We find that the model performance remains mostly unaffected, while the model shows much improved interpretability. Additionally, we verify the interpretation of the bottleneck components with an intervention experiment using activation patching.

时间序列可解释性Transformer概念瓶颈

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。