数据去相关能改善模型解释的可靠性,但效果因方法而异。
The effect of whitening on explanation performance
- 用五种去相关技术处理数据,测试16种解释方法的表现。
- 部分方法在去相关后解释准确率显著提升,但提升幅度差异大。
- 适合关注模型可解释性与预处理影响的研究者参考。
可解释人工智能(XAI)旨在为机器学习模型提供透明洞察,但许多特征归因方法的可靠性仍存疑。已有研究(Haufe等,2014;Wilming等,2022, 2023)表明,这些方法常错误地赋予非信息变量(如抑制变量)高重要性,导致根本性误读。由于统计抑制由特征依赖引起,本研究探讨数据白化(一种常见去相关预处理)是否可缓解此类错误。基于已建立的XAI-TRIS基准(Clark等,2024b),该基准提供合成真实数据和解释正确性的量化指标,我们实证评估了16种主流特征归因方法在结合5种不同白化变换下的表现。此外,我们分析了一个最小线性二维分类问题(Wilming等,2023),从理论上评估白化能否消除贝叶斯最优模型中抑制特征的影响。结果表明,尽管特定白化技术可提升解释性能,但改善程度在不同XAI方法和模型架构间差异显著。这些发现凸显了数据非线性、预处理质量与归因保真度之间的复杂关系,强调预处理在提升模型可解释性中的关键作用。
原文摘要 · Abstract (English)
Explainable Artificial Intelligence (XAI) aims to provide transparent insights into machine learning models, yet the reliability of many feature attribution methods remains a critical challenge. Prior research (Haufe et al., 2014; Wilming et al., 2022, 2023) has demonstrated that these methods often erroneously assign significant importance to non-informative variables, such as suppressor variables, leading to fundamental misinterpretations. Since statistical suppression is induced by feature dependencies, this study investigates whether data whitening, a common preprocessing technique for decorrelation, can mitigate such errors. Using the established XAI-TRIS benchmark (Clark et al., 2024b), which offers synthetic ground-truth data and quantitative measures of explanation correctness, we empirically evaluate 16 popular feature attribution methods applied in combination with 5 distinct whitening transforms. Additionally, we analyze a minimal linear two-dimensional classification problem (Wilming et al., 2023) to theoretically assess whether whitening can remove the impact of suppressor features from Bayes-optimal models. Our results indicate that, while specific whitening techniques can improve explanation performance, the degree of improvement varies substantially across XAI methods and model architectures. These findings highlight the complex relationship between data non-linearities, preprocessing quality, and attribution fidelity, underscoring the vital role of pre-processing techniques in enhancing model interpretability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。