arXiv:2603.15821cs.LGcs.AI2026-03

不同模型结构导致解释差异,即使预测结果相同。

Hypothesis Class Determines Explanation: Why Accurate Models Disagree on Feature Attribution

  • 同一数据上训练的模型,若结构不同则解释差异大。
  • 同结构模型解释高度一致,跨结构模型一致性接近随机水平。
  • 提出解释可靠性评分,可判断解释是否稳定且无需重训。

解释性AI中普遍假设预测能力相同的模型应产生相似解释,但本研究通过24个数据集的大规模实证发现,这一假设不成立。相同预测性能的模型可能产生显著不同的特征归因。这种分歧具有结构性:同假设类模型间解释高度一致,而跨类模型(如树模型与线性模型)在相同数据划分下一致性明显下降,通常处于或低于彩票阈值。我们识别出假设类是该现象的根本原因,称之为“解释彩票”。理论上证明,在数据生成过程存在交互结构时,该一致性差距仍存在。据此提出后验诊断指标解释可靠性分数R(x),可预测解释在不同架构间的稳定性,无需额外训练。结果表明,模型选择并非解释中立:部署时所选假设类将决定哪些特征被赋予决策责任。

原文摘要 · Abstract (English)

The assumption that prediction-equivalent models produce equivalent explanations underlies many practices in explainable AI, including model selection, auditing, and regulatory evaluation. In this work, we show that this assumption does not hold. Through a large-scale empirical study across 24 datasets and multiple model classes, we find that models with identical predictive behavior can produce substantially different feature attributions. This disagreement is highly structured: models within the same hypothesis class exhibit strong agreement, while cross-class pairs (e.g., tree-based vs. linear) trained on identical data splits show substantially reduced agreement, consistently near or below the lottery threshold. We identify hypothesis class as the structural driver of this phenomenon, which we term the Explanation Lottery. We theoretically show that the resulting Agreement Gap persists under interaction structure in the data-generating process. This structural finding motivates a post-hoc diagnostic, the Explanation Reliability Score R(x), which predicts when explanations are stable across architectures without additional training. Our results demonstrate that model selection is not explanation-neutral: the hypothesis class chosen for deployment can determine which features are attributed responsibility for a decision.

可解释性特征归因模型选择解释稳定性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。