提出可识别的时间序列归因图方法,提升模型解释可靠性。
Time-series attribution maps with regularized contrastive learning
- 用正则化对比学习+反向神经梯度生成归因图
- 在合成数据上准确识别真实归因图的零与非零项
- 适合研究神经动力学和模型决策过程的科研人员
基于梯度的归因方法旨在解释深度学习模型的决策,但迄今缺乏可识别性保证。本文提出一种新方法,通过在时间序列数据上训练正则化对比学习算法,并结合新型归因方法Inverted Neuron Gradient(统称xCEBRA),生成具有可识别性保证的归因图。理论上证明xCEBRA能有效识别数据生成过程的雅可比矩阵;实验上,在合成数据集上实现对真实归因图中零值与非零值项的稳健逼近,并显著优于基于特征消融、Shapley值及其他梯度方法的现有方法。本工作首次实现了时间序列归因图的可识别推断,为理解神经动力学及神经网络内部决策机制开辟新路径。
原文摘要 · Abstract (English)
Gradient-based attribution methods aim to explain decisions of deep learning models but so far lack identifiability guarantees. Here, we propose a method to generate attribution maps with identifiability guarantees by developing a regularized contrastive learning algorithm trained on time-series data plus a new attribution method called Inverted Neuron Gradient (collectively named xCEBRA). We show theoretically that xCEBRA has favorable properties for identifying the Jacobian matrix of the data generating process. Empirically, we demonstrate robust approximation of zero vs. non-zero entries in the ground-truth attribution map on synthetic datasets, and significant improvements across previous attribution methods based on feature ablation, Shapley values, and other gradient-based methods. Our work constitutes a first example of identifiable inference of time-series attribution maps and opens avenues to a better understanding of time-series data, such as for neural dynamics and decision-processes within neural networks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。