用变换测试法评估模型解释的可靠性,无需真实标签
Metamorphic Testing with the Rashomon Set: Explanation Faithfulness in Machine Learning
- 通过变换测试检验解释与模型行为的一致性
- 在4个数据集上验证了不同解释方法的可信度差异
- 适合关注模型解释可信性的研究人员使用
多个机器学习模型可在同一任务上达到相近的预测性能,但提供的基于特征的解释却大相径庭,这种现象被称为可解释机器学习中的Rashomon效应,引发对哪些解释具有可信度的疑问。我们提出一种基于变换测试的框架,无需真实标签即可评估解释的忠实性,通过探索后验解释方法产生的特征重要性来实现。该框架定义了五种变换关系,用于形式化模型行为与特征归因之间的预期一致性。我们在两个表格回归数据集和两种后验解释器(SHAP与LIME)上应用该方法,展示了其有效性。该框架提供了一种实用、模型无关的工具,可用于筛选具备可靠且可信解释的准确模型。
原文摘要 · Abstract (English)
Multiple machine learning models can achieve near-equivalent predictive performance on the same task, yet provide divergent feature-based explanations. This is called the Rashomon effect of (explainable) machine learning, and it raises the question of which explanations, if any, are trustworthy. We propose a framework based on metamorphic testing that assesses explanation faithfulness without requiring ground-truth labels by exploring attributed feature importance from post-hoc explanation methods. Five metamorphic relations formalize expected consistency properties between model behavior and feature attributions. We apply this general framework to two tabular regression datasets and two post-hoc explainers (SHAP and LIME) to demonstrate the approach. The framework offers a practical, model-agnostic tool for selecting accurate models with reliable and trustworthy explanations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。