arXiv:2601.12654cs.LGcs.AI2026-01被引 8

SHAP解释结果不一致,同一预测多次运行得不同结论。

Explanation Multiplicity in SHAP: Characterization and Assessment

  • 提出解释多重性概念,分析SHAP结果波动的根源。
  • 发现即使高置信度预测,解释仍频繁变化。
  • 建议用随机基线对比评估解释稳定性,适合监管与审计者。

后验解释广泛用于高风险领域(如信贷、招聘、医疗)中解释自动化决策。其中,SHAP常被视为可靠地识别影响单个预测的关键特征,并被用于支持补救、监督与问责。然而实践中,即使个体、任务和模型固定,重复运行仍可能导致显著不同的SHAP解释。本文将此现象称为‘解释多重性’:对同一决策存在多个内部合理但实质不同的解释。这挑战了负责任AI部署的规范基础,削弱了解释可追溯至不利结果原因的预期。本文提出系统方法,刻画后验特征归因方法中的解释多重性,区分由模型训练与选择引发的差异,以及解释流程本身固有的随机性。此外,是否察觉到多重性取决于评估一致性的方式;常用幅度指标可能显示稳定,却掩盖了重要特征身份与排序的剧烈变动。为提供上下文参照,我们推导并估计在合理零模型下的随机基线值,形成解释分歧的规范参考点。在多个数据集、模型类别和置信区间下,均发现解释多重性普遍存在,且在高度受控条件下(如高置信度预测)依然持续。因此,解释实践必须采用与其社会角色一致的评估指标与基线。

原文摘要 · Abstract (English)

Post-hoc explanations are widely used to justify, contest, and review automated decisions in high-stakes domains such as lending, employment, and healthcare. Among these methods, SHAP is often treated as providing a reliable account of which features mattered for an individual prediction and is routinely used to support recourse, oversight, and accountability. In practice, however, SHAP explanations can differ substantially across repeated runs, even when the individual, prediction task, and trained model are held fixed. We conceptualize and name this phenomenon explanation multiplicity: the existence of multiple, internally valid but substantively different explanations for the same decision. Explanation multiplicity poses a normative challenge for responsible AI deployment, as it undermines expectations that explanations can reliably identify the reasons for an adverse outcome. We present a comprehensive methodology for characterizing explanation multiplicity in post-hoc feature attribution methods, disentangling sources arising from model training and selection versus stochasticity intrinsic to the explanation pipeline. Furthermore, whether explanation multiplicity is surfaced depends on how explanation consistency is measured. Commonly used magnitude-based metrics can suggest stability while masking substantial instability in the identity and ordering of top-ranked features. To contextualize observed instability, we derive and estimate randomized baseline values under plausible null models, providing a principled reference point for interpreting explanation disagreement. Across datasets, model classes, and confidence regimes, we find that explanation multiplicity is widespread and persists even under highly controlled conditions, including high-confidence predictions. Thus explanation practices must be evaluated using metrics and baselines aligned with their intended societal role.

SHAP解释可信度多重性可解释AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。