医生用AI诊断胸片时,文字解释易导致过度依赖,图文结合更安全。
Fool Me Once? Contrasting Textual and Visual Explanations in a Clinical Decision-Support Setting
- 对比视觉热力图、文字说明及两者结合的效果
- 文字解释使医生误信错误AI建议,图文结合降低依赖
- 解释内容真实且与AI判断一致时,才真正有用
随着AI模型能力提升,其在安全关键领域应用日益广泛。可解释AI(XAI)旨在通过透明化推理过程提升使用安全性。然而,现有解释方法极少在真实用户场景中评估。为此,我们在人机协作胸片分析场景下,对85名医疗从业者开展了大规模用户研究,评估了三类解释:视觉热力图、自然语言说明及二者结合。重点考察不同解释类型在AI建议与解释本身是否正确时对用户的影响。结果发现,文字解释导致显著的过度依赖,而结合热力图可缓解此问题。此外,解释质量——即信息的真实性及其与AI判断的一致性——显著影响各类解释的实际效用。
原文摘要 · Abstract (English)
The growing capabilities of AI models are leading to their wider use, including in safety-critical domains. Explainable AI (XAI) aims to make these models safer to use by making their inference process more transparent. However, current explainability methods are seldom evaluated in the way they are intended to be used: by real-world end users. To address this, we conducted a large-scale user study with 85 healthcare practitioners in the context of human-AI collaborative chest X-ray analysis. We evaluated three types of explanations: visual explanations (saliency maps), natural language explanations, and a combination of both modalities. We specifically examined how different explanation types influence users depending on whether the AI advice and explanations are factually correct. We find that text-based explanations lead to significant over-reliance, which is alleviated by combining them with saliency maps. We also observe that the quality of explanations, that is, how much factually correct information they entail, and how much this aligns with AI correctness, significantly impacts the usefulness of the different explanation types.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。