测试算法评价指标是否符合用户对反事实解释的感知,发现两者关联很弱。
Do Metrics for Counterfactual Explanations Align with User Perception?
- 通过真人评分对比算法指标与用户感知
- 多数指标与用户评价相关性低且依赖数据集
- 增加指标数量无法提升预测准确率,暴露现有方法缺陷
可解释性被视为可信人工智能系统的关键。然而,当前用于评估反事实解释的指标多为算法指标,很少经过人类判断的验证。本研究通过实证分析,直接比较三种数据集上算法指标与人类对解释质量的多维度评价。参与者对反事实解释进行评分,我们将其与一系列标准指标进行关联分析。结果表明,算法指标与人类评分的相关性普遍较弱,且高度依赖数据集;增加指标数量也未带来预测性能的可靠提升,反映出现有指标在捕捉人类关注点上的结构性局限。研究指出,当前广泛使用的反事实评估指标未能反映用户感知的核心质量维度,亟需更以人为中心的可解释人工智能评估方法。
原文摘要 · Abstract (English)
Explainability is widely regarded as essential for trustworthy artificial intelligence systems. However, the metrics commonly used to evaluate counterfactual explanations are algorithmic evaluation metrics that are rarely validated against human judgments of explanation quality. This raises the question of whether such metrics meaningfully reflect user perceptions. We address this question through an empirical study that directly compares algorithmic evaluation metrics with human judgments across three datasets. Participants rated counterfactual explanations along multiple dimensions of perceived quality, which we relate to a comprehensive set of standard counterfactual metrics. We analyze both individual relationships and the extent to which combinations of metrics can predict human assessments. Our results show that correlations between algorithmic metrics and human ratings are generally weak and strongly dataset-dependent. Moreover, increasing the number of metrics used in predictive models does not lead to reliable improvements, indicating structural limitations in how current metrics capture criteria relevant for humans. Overall, our findings suggest that widely used counterfactual evaluation metrics fail to reflect key aspects of explanation quality as perceived by users, underscoring the need for more human-centered approaches to evaluating explainable artificial intelligence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。