arXiv:2506.12176cs.LGstat.ML2025-06中稿 · the Workshop on Sc…

高保真解释未必揭示数据本质,可能误导对模型性能的理解

"Faithful to What?" On the Limits of Fidelity-Based Explanations

  • 用线性可解度评分衡量神经网络输入输出的线性特征
  • 高保真代理模型在多个数据集上仍无法还原任务关键结构
  • 适合关注解释可靠性与模型性能差异的研究者

在可解释AI中,代理模型通常通过其与神经网络预测的一致性(保真度)来评估。然而,保真度衡量的是与学习模型的一致性,而非与任务背后数据生成信号的一致性。本文提出线性评分λ(f),用于量化回归网络输入-输出行为的线性可解程度。λ(f)定义为代理模型对网络预测的拟合优度R²。在合成和真实回归数据集上,我们发现代理模型即使具有高保真度,也无法恢复区分神经网络与简单模型的预测优势。在多个案例中,高保真代理模型甚至劣于直接在数据上训练的线性基线模型。结果表明,解释模型行为不等于解释数据中的任务相关结构,揭示了基于保真度的解释在推断预测性能时的局限性。

原文摘要 · Abstract (English)

In explainable AI, surrogate models are commonly evaluated by their fidelity to a neural network's predictions. Fidelity, however, measures alignment to a learned model rather than alignment to the data-generating signal underlying the task. This work introduces the linearity score $λ(f)$, a diagnostic that quantifies the extent to which a regression network's input--output behavior is linearly decodable. $λ(f)$ is defined as an $R^2$ measure of surrogate fit to the network. Across synthetic and real-world regression datasets, we find that surrogates can achieve high fidelity to a neural network while failing to recover the predictive gains that distinguish the network from simpler models. In several cases, high-fidelity surrogates underperform even linear baselines trained directly on the data. These results demonstrate that explaining a model's behavior is not equivalent to explaining the task-relevant structure of the data, highlighting a limitation of fidelity-based explanations when used to reason about predictive performance.

可解释AI模型解释保真度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。