arXiv:2602.09238cs.LG2026-02被引 3

模型解释的权重更多来自图像显著性,而非真实信息关联。

Feature salience - not task-informativeness - drives machine learning model explanations

  • 用透明水印测试不同场景,发现解释结果受图像显著性主导
  • 水印区域重要性评分普遍高于45%(R²≥0.45),与任务信息无关
  • 适合关注模型解释可信度的研究者和开发者

可解释人工智能(XAI)旨在揭示机器学习模型的决策过程,其中一个重要目标是识别如捷径学习等缺陷。这一承诺依赖于一个假设:被XAI标记为重要的输入特征必须包含目标变量的信息。然而,实际中重要性归因是否主要由信息量驱动,还是受统计抑制、测试时新奇性或高特征显著性等数据属性影响尚不明确。为此,我们在三个变体的二分类图像任务上训练深度学习模型,其中透明水印分别缺失、作为类别相关混杂因素或作为类别无关噪声。五种主流归因方法的结果显示,无论训练设置如何,水印区域的相对重要性均显著升高(R² ≥ 0.45)。相比之下,水印是否与类别相关对相对重要性的影响微乎其微(R² ≤ 0.03),尽管这显著影响模型性能和泛化能力。XAI方法的行为类似模型无关的边缘检测滤波器,当图像亮度以较小特征值编码时,对水印的重要性评估明显降低。这些结果表明,重要性归因最强烈地由测试时图像结构的显著性驱动,而非模型所学的统计关联。此前展示成功XAI应用的研究应重新审视特征显著性与信息量偶然共现可能带来的误导,使用特征归因方法作为构建模块的工作流也需严格审查。

原文摘要 · Abstract (English)

Explainable AI (XAI) promises to provide insight into machine learning models' decision processes, where one goal is to identify failures such as shortcut learning. This promise relies on the field's assumption that input features marked as important by an XAI must contain information about the target variable. However, it is unclear whether informativeness is indeed the main driver of importance attribution in practice, or if other data properties such as statistical suppression, novelty at test-time, or high feature salience substantially contribute. To clarify this, we trained deep learning models on three variants of a binary image classification task, in which translucent watermarks are either absent, act as class-dependent confounds, or represent class-independent noise. Results for five popular attribution methods show substantially elevated relative importance in watermarked areas (RIW) for all models regardless of the training setting ($R^2 \geq .45$). By contrast, whether the presence of watermarks is class-dependent or not only has a marginal effect on RIW ($R^2 \leq .03$), despite a clear impact impact on model performance and generalisation ability. XAI methods show similar behaviour to model-agnostic edge detection filters and attribute substantially less importance to watermarks when bright image intensities are encoded by smaller instead of larger feature values. These results indicate that importance attribution is most strongly driven by the salience of image structures at test time rather than statistical associations learned by machine learning models. Previous studies demonstrating successful XAI application should be reevaluated with respect to a possibly spurious concurrency of feature salience and informativeness, and workflows using feature attribution methods as building blocks should be scrutinised.

可解释AI特征重要性模型可信度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。