用反向解释对比,更精准评估模型解释质量。
Evaluate with the Inverse: Efficient Approximation of Latent Explanation Quality Distribution
- 提出质量差距估计法,以反向解释为参照比较
- 在多个模型和数据集上验证优于随机对照方法
- 提升评估可靠性,适合模型审计与对齐研究
高质量模型解释有助于开发者识别偏差、对齐人类价值观并确保伦理合规。可解释人工智能(XAI)从业者依赖特定度量来评估解释质量,如忠实性、定位精度和鲁棒性。然而这些度量未能解决关键问题:当前解释质量如何与其他可能解释相比?传统方法通过与随机生成解释对比进行评估。本文提出质量差距估计(QGE),直接与概念上的“反向解释”——即原解释的对立面——进行比较。在多种模型架构、数据集和成熟度量下的大量测试表明,QGE显著优于传统方法,并提升了评估的统计可靠性。这一进展推动了对模型行为更深入的评估,助力更有效的模型审查。
原文摘要 · Abstract (English)
Obtaining high-quality explanations of a model's output enables developers to identify and correct biases, align the system's behavior with human values, and ensure ethical compliance. Explainable Artificial Intelligence (XAI) practitioners rely on specific measures to gauge the quality of such explanations. These measures assess key attributes, such as how closely an explanation aligns with a model's decision process (faithfulness), how accurately it pinpoints the relevant input features (localization), and its consistency across different cases (robustness). Despite providing valuable information, these measures do not fully address a critical practitioner's concern: how does the quality of a given explanation compare to other potential explanations? Traditionally, the quality of an explanation has been assessed by comparing it to a randomly generated counterpart. This paper introduces an alternative: the Quality Gap Estimate (QGE). The QGE method offers a direct comparison to what can be viewed as the `inverse' explanation, one that conceptually represents the antithesis of the original explanation. Our extensive testing across multiple model architectures, datasets, and established quality metrics demonstrates that the QGE method is superior to the traditional approach. Furthermore, we show that QGE enhances the statistical reliability of these quality assessments. This advance represents a significant step toward a more insightful evaluation of explanations that enables a more effective inspection of a model's behavior.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。