arXiv:2507.15162cs.LG2025-07被引 1

提出用户中心的解释评估模型,让AI解释更符合真实用户需求。

Designing User-Centric Metrics for Evaluation of Counterfactual Explanations

  • 通过用户实验发现现有评估指标与用户偏好不符
  • 提出AWP模型,准确预测用户偏好的解释达84.37%
  • 适合关注可解释AI和用户体验的研究者

反事实解释(CFEs)因其能提供可操作建议而受到欢迎,即识别使模型预测改变为更理想结果所需的最小特征值调整。然而,大多数先前研究依赖人工设计的评估指标(如接近度),可能忽视终端用户的实际偏好与约束,例如用户对修改某些特征所需努力的感知可能不同于模型设计者。为填补这一空白,本文有三项创新贡献:首先,我们通过对20名亚马逊MTurk众包工作者的试点研究,验证了现有评估指标与真实用户偏好的对齐程度。结果显示,用户偏好的反事实解释仅在63.81%的情况下与基于接近度的解释一致,表明这些指标在真实场景中的适用性有限。其次,受此启发,我们开展为期两天的用户研究,共41名参与者在真实的信贷申请场景中评估反事实解释,以检验三种关于用户如何评估反事实解释的直觉假设。第三,基于该研究结果,我们提出了一个用户中心的两阶段模型——AWP,描述了用户评估和选择反事实解释的一种可能机制。实验表明,该模型能以84.37%的准确率预测用户偏好的解释。本研究首次提供了针对个性化成本模型在反事实生成中的人类中心验证,强调了开发自适应、用户中心评估指标的重要性。

原文摘要 · Abstract (English)

Counterfactual Explanations (CFEs) have grown in popularity as a means of offering actionable guidance by identifying the minimum changes in feature values required to flip an ML model's prediction to something more desirable. Unfortunately, most prior research on CFEs relies on artificial evaluation metrics, such as proximity, which may overlook end-user preferences and constraints, e.g., the user's perception of effort needed to make certain feature changes may differ from that of the model designer. To address this research gap, this paper makes three novel contributions. First, we conduct a pilot study with 20 crowd-workers on Amazon MTurk to experimentally validate the alignment of existing CF evaluation metrics with real-world user preferences. Results show that user-preferred CFEs matched those based on proximity in only 63.81% of cases, highlighting the limited applicability of these metrics in real-world settings. Second, inspired by the need to design a user-informed evaluation metric for CFEs, we conduct a more detailed two-day user study with 41 participants facing realistic credit application scenarios to find experimental support for or against three intuitive hypotheses that may explain how end users evaluate CFEs. Third, based on the findings of this second study, we propose the AWP model, a novel user-centric, two-stage model that describes one possible mechanism by which users evaluate and select CFEs. Our results show that AWP predicts user-preferred CFEs with 84.37% accuracy. Our study provides the first human-centered validation for personalized cost models in CFE generation and highlights the need for adaptive, user-centered evaluation metrics.

可解释AI用户研究反事实解释

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。