arXiv:2602.02763cs.LG2026-02中稿 · ICML被引 1

提出双目标攻击,让时间序列分类器的预测与解释可被分别操控。

Exposing Vulnerabilities in Explanation for Time Series Classifiers via Dual-Target Attacks

  • 设计双目标攻击TSEF,同时操纵分类器输出和解释结果
  • 攻击后预测错误但解释仍符合指定参考依据,保持一致性
  • 揭示现有解释稳定性评估不可靠,适合可信AI研究者参考

可解释的时间序列深度学习系统常通过检查解释的时间一致性来评估鲁棒性,隐含假设此即为可靠证据。我们发现该假设可能失效:预测与解释可被对抗性地解耦,实现目标误分类的同时,解释仍保持合理且与选定参考理由一致。为此提出TSEF(时间序列解释欺骗者)——一种联合操控分类器与解释器输出的双目标攻击。与仅破坏解释的单目标误分类攻击不同,TSEF在实现目标预测变更的同时,使解释保持与参考一致。在多个数据集与解释器骨干网络上,结果均显示解释稳定性是决策鲁棒性的误导性代理,呼吁在可信时间序列任务中引入耦合感知的鲁棒性评估。

原文摘要 · Abstract (English)

Interpretable time series deep learning systems are often assessed by checking temporal consistency on explanations, implicitly treating this as evidence of robustness. We show that this assumption can fail: Predictions and explanations can be adversarially decoupled, enabling targeted misclassification while the explanation remains plausible and consistent with a chosen reference rationale. We propose TSEF (Time Series Explanation Fooler), a dual-target attack that jointly manipulates the classifier and explainer outputs. In contrast to single-objective misclassification attacks that disrupt explanation and spread attribution mass broadly, TSEF achieves targeted prediction changes while keeping explanations consistent with the reference. Across multiple datasets and explainer backbones, our results consistently reveal that explanation stability is a misleading proxy for decision robustness and motivate coupling-aware robustness evaluations for trustworthy time series tasks.

时间序列对抗攻击可解释性鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。