arXiv:2608.25897cs.LGcs.AI2026-08

统一时间序列解释框架,解决归因与反事实解释的不稳定性问题

Towards A Unified Information Bottleneck Framework for Time Series Explanations

论文配图:Towards A Unified Information Bottleneck Framework for Time Series Explanations
图 1 · 摘自论文原文
  • 基于信息瓶颈原理,构建统一解释目标函数
  • 在合成与真实数据上均优于现有方法,解释更忠实稳定
  • 适合需要可解释性与鲁棒性的时序模型应用

解释作用于时间序列数据的深度学习模型至关重要,尤其在需要透明决策的场景中。现有解释方法主要分为两类:基于归因的解释识别预测关键时间片段,以及基于反事实的解释揭示输入如何修改才能改变模型判断。然而这两类方法长期独立发展,导致归因缺乏因果验证,反事实解释易生成类似对抗噪声的不稳定结果。本文从信息论视角重新审视时间序列可解释性,指出现有方法对平凡解和分布偏移敏感。为此,我们提出一个统一的目标函数,将归因与反事实推理整合到单一框架中。基于信息瓶颈原理,该函数显式避免平凡解释和域外反事实。据此,我们设计了 {\modelname},一种通过参数化变换网络生成嵌入解释实例的新框架:保留的信息用于归因解释,可控信息移除则生成稳定反事实解释。我们在合成与真实世界基准上对 {\modelname} 进行评估,结果表明其在定量与定性指标上均显著优于主流基线方法,实现高保真归因与稳定反事实。

原文摘要 · Abstract (English)

Explaining deep learning models operating on time series data is crucial in various applications that require transparent and interpretable insights into model behavior. {Existing explanation methods generally fall into two categories: attribution-based explanations, which identify the temporal regions most responsible for a prediction, and counterfactual explanations, which reveal how an input should be modified to alter the model's decision.} {Despite valuable insights, these two fields are largely studied independently. This disconnect leaves attribution methods lacking causal validation, while counterfactual methods suffer from severe instability, producing adversarial-like noise instead of meaningful explanations.} In this work, we revisit time-series explainability from an information-theoretic perspective and show that existing explainers are vulnerable to trivial solutions and distributional shifts. To address these limitations, we propose a unified objective function for explainable time series learning that bridges attribution and counterfactual reasoning within a single framework. Building upon the Information Bottleneck principle, our formulation explicitly prevents trivial explanations and out-of-distribution counterfactuals. {Based on this objective function, we introduce {\modelname}, a novel explanation framework that learns a parametric transformation network to construct explanation-embedded instances, where preserved information yields attribution explanations and controlled information removal produces stable counterfactual explanations.} We evaluate {\modelname} on synthetic and real-world benchmarks against state-of-the-art baselines. Extensive quantitative and qualitative results show that {\modelname} consistently outperforms competing methods, yielding faithful attributions and stable counterfactual explanations.

时间序列可解释性信息瓶颈反事实

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。