研究发现,AI解释对视觉任务中的学习效果有限。
Confident Teacher, Confident Student? A Novel User Study Design for Investigating the Didactic Potential of Explanations and their Impact on Uncertainty
- 设计1200人用户实验,评估AI解释在生物物种标注中的作用
- 有解释时准确率提升,但仅看预测结果效果相近
- 解释反而让用户更盲从模型错误,且无长期学习收益
评估可解释人工智能(XAI)中解释质量至今仍是难题,学界存在争议。本文提出一种新型实验设计,评估XAI在人机协作与教学中的潜力。通过1200名参与者在复杂生物分类标注任务上的研究,发现使用AI辅助可提升标注准确率并降低不确定性。然而,仅展示模型预测与同时提供解释相比,准确率提升并无显著差异。此外,解释带来负面影响:用户更倾向于复制模型的错误预测。在评估教学效果时,接受过AI协助的用户后续标注表现未显著改善,表明解释在视觉人机协作中难以产生持久学习效应。所有代码与数据详见GitHub:https://github.com/TeodorChiaburu/beexplainable。
原文摘要 · Abstract (English)
Evaluating the quality of explanations in Explainable Artificial Intelligence (XAI) is to this day a challenging problem, with ongoing debate in the research community. While some advocate for establishing standardized offline metrics, others emphasize the importance of human-in-the-loop (HIL) evaluation. Here we propose an experimental design to evaluate the potential of XAI in human-AI collaborative settings as well as the potential of XAI for didactics. In a user study with 1200 participants we investigate the impact of explanations on human performance on a challenging visual task - annotation of biological species in complex taxonomies. Our results demonstrate the potential of XAI in complex visual annotation tasks: users become more accurate in their annotations and demonstrate less uncertainty with AI assistance. The increase in accuracy was, however, not significantly different when users were shown the mere prediction of the model compared to when also providing an explanation. We also find negative effects of explanations: users tend to replicate the model's predictions more often when shown explanations, even when those predictions are wrong. When evaluating the didactic effects of explanations in collaborative human-AI settings, we find that users' annotations are not significantly better after performing annotation with AI assistance. This suggests that explanations in visual human-AI collaboration do not appear to induce lasting learning effects. All code and experimental data can be found in our GitHub repository: https://github.com/TeodorChiaburu/beexplainable.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。