解释稳定性取决于模型与方法的组合,而非模型本身。
Position: Explanation Stability Is a Property of the Model Method Pair, Not the Model

- 提出跨方法验证解释稳定性,避免单一方法误导
- 三种模型在不同方法下稳定性排名反转,最大差异达17.3%
- 适合关注AI可解释性可信度的研究者和监管方
本文主张,若无跨方法验证,关于解释稳定性的说法在科学上无效。如同统计显著性需明确定义检验统计量,稳定性也应基于多个归因范式评估,或明确限定于单个方法的计算目标。在受控的胸部X光实验中,DenseNet201、ResNet50V2和InceptionV3的AUC均超过99%,但其稳定性排序在不同归因方法间反转。LayerCAM将InceptionV3列为最稳定模型(IoU=0.777),而GradCAM++则更偏好DenseNet201,使InceptionV3的稳定性得分下降17.3%。结果表明,解释稳定性是模型与方法配对的涌现属性,而非模型固有特征。因此,我们主张解释性结论应跨方法验证,且监管提交需明确说明所用归因算子,以避免制造虚假的安全保障。
原文摘要 · Abstract (English)
This position paper argues that claims about explanation stability are scientifically invalid without cross method validation. Just as statistical significance requires the test statistic to be specified, stability should either be evaluated across multiple attribution paradigms or explicitly scoped to the computational objective of a single method. In controlled chest X ray experiments, DenseNet201, ResNet50V2, and InceptionV3 achieved AUC values above 99%, yet their stability rankings reversed across attribution methods. LayerCAM ranked InceptionV3 as the most stable model, with an IoU of 0.777, whereas GradCAM++ favored DenseNet201 and reduced InceptionV3 stability score by 17.3%. These findings demonstrate that explanation stability is an emergent property of the model method pair rather than an intrinsic characteristic of the model alone. We therefore argue that explanation based claims should be validated across multiple attribution methods and that regulatory submissions should explicitly specify the attribution operators used to avoid creating illusory safety assurances.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。