区间型反事实解释更有效提升用户对AI的理解与信任。
Improving understanding and trust in AI: How users benefit from interval-based counterfactual explanations

- 用区间而非单点提供反事实解释,效果更佳。
- 区间解释使用户理解力和信任度显著提升。
- 不同认知风格者对解释的反应差异明显。
现有研究极少评估黑箱模型后验解释的类型对用户的影响。本研究通过在线用户实验,比较了四种解释方式:无解释(对照组)、特征重要性评分、点式反事实解释和区间式反事实解释。实验采用被试内设计,结果显示,相较于其他解释类型,区间式解释在提升用户对模型的理解和表现出的信任方面具有明显优势。研究未验证此前部分研究关于点式反事实解释优于对照组的结论。此外,个体差异如认知风格或人格特质显著影响解释效果。
原文摘要 · Abstract (English)
Experimental user studies evaluating the effectiveness of different subtypes of post-hoc explanations for black-box models are largely nonexistent. Therefore, the aim of this study was to investigate and evaluate how different types of counterfactual explanations, namely single point explanations and interval-based explanations, affect both model understanding and (demonstrated) trust. We conducted an online user study using a within-subjects experimental design, where the experimental arms were (i) no explanation (control), (ii) feature importance scores, (iii) point counterfactual explanations, and (iv) interval counterfactual explanations. Our results clearly show the superiority of interval explanations over other tested explanation types in increasing both model understanding and demonstrated trust in the AI. We could not support findings of some previous studies showing an effect of point counterfactual explanations compared to the control group. Our results further highlight the role individual differences in, for example, cognitive style or personality, in explanation effectiveness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。