arXiv:2505.11210cs.LGstat.ML2025-05NeurIPS被引 3

解决非线性模型解释中的假阳性误判问题,提升解释可靠性。

Minimizing False-Positive Attributions in Explanations of Non-Linear Models

  • 通过局部线性代理生成可解释的生成表示,抑制干扰变量影响。
  • 在XAI-TRIS基准上优于其他方法,显著降低假阳性归因率。
  • 适用于需要可靠解释的场景,如脑电医疗分析。

抑制变量可能在不依赖目标结果的情况下影响模型预测,对可解释人工智能(XAI)方法构成重大挑战。这些变量可能导致特征归因的假阳性,削弱解释的实用性。尽管线性模型已有有效解决方案,但其向非线性模型和基于实例的解释扩展仍受限。我们提出PatternLocal,一种新型XAI技术,填补这一空白。PatternLocal以局部线性代理(如LIME、KernelSHAP或基于梯度的方法)为基础,将判别模型权重转化为生成表示,从而抑制抑制变量的影响,同时保持局部保真度。在XAI-TRIS基准上进行大规模超参数优化后,PatternLocal持续优于其他XAI方法,并在解释非线性任务时减少了假阳性归因,实现更可靠、可操作的洞察。我们进一步在EEG运动想象数据集上评估PatternLocal,验证了其具有生理学合理性的解释。

原文摘要 · Abstract (English)

Suppressor variables can influence model predictions without being dependent on the target outcome, and they pose a significant challenge for Explainable AI (XAI) methods. These variables may cause false-positive feature attributions, undermining the utility of explanations. Although effective remedies exist for linear models, their extension to non-linear models and instance-based explanations has remained limited. We introduce PatternLocal, a novel XAI technique that addresses this gap. PatternLocal begins with a locally linear surrogate, e.g., LIME, KernelSHAP, or gradient-based methods, and transforms the resulting discriminative model weights into a generative representation, thereby suppressing the influence of suppressor variables while preserving local fidelity. In extensive hyperparameter optimization on the XAI-TRIS benchmark, PatternLocal consistently outperformed other XAI methods and reduced false-positive attributions when explaining non-linear tasks, thereby enabling more reliable and actionable insights. We further evaluate PatternLocal on an EEG motor imagery dataset, demonstrating physiologically plausible explanations.

可解释AI特征归因非线性模型脑电信号

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。