系统调研临床决策系统中可解释AI的人因评估,发现方法局限与认知负担问题。
A Survey on Human-Centered Evaluation of Explainable AI Methods in Clinical Decision Support Systems
- 基于PRISMA框架分析31项人因实验,聚焦后处理、模型无关方法如SHAP和Grad-CAM
- 超80%研究采用此类方法,但临床医生样本量普遍低于25人,结果易受偏差影响
- 解释虽提升信任感,却常加重认知负荷,建议构建以利益相关者为中心的评估框架
可解释人工智能(XAI)对临床决策支持系统(CDSS)的透明性与实际应用至关重要。然而,现有XAI方法的真实效果有限且评估不一致。本研究基于PRISMA指南,系统调研了31项应用于CDSS的可解释AI人因评估(HCE),按XAI方法、评估设计及采纳障碍进行分类。结果显示,多数研究采用后处理、模型无关方法(如SHAP、Grad-CAM),通常通过小规模医生实验验证。超过80%的研究使用此类方法,且医生样本量普遍低于25人。结果表明,解释能提升医生信任度与诊断信心,但常导致认知负荷增加,并与临床推理过程存在错位。为弥合此差距,我们提出一种以利益相关者为中心的评估框架,融合社会技术原则与人机交互理念,以指导未来可信且临床可行的XAI-CDSS发展。
原文摘要 · Abstract (English)
Explainable Artificial Intelligence (XAI) is essential for the transparency and clinical adoption of Clinical Decision Support Systems (CDSS). However, the real-world effectiveness of existing XAI methods remains limited and is inconsistently evaluated. This study conducts a systematic PRISMA-guided survey of 31 human-centered evaluations (HCE) of XAI applied to CDSS, classifying them by XAI methodology, evaluation design, and adoption barrier. Our findings reveal that most existing studies employ post-hoc, model-agnostic approaches such as SHAP and Grad-CAM, typically assessed through small-scale clinician studies. The results show that over 80% of the studies adopt post-hoc, model-agnostic approaches such as SHAP and Grad-CAM, and that clinician sample sizes remain below 25 participants. The findings indicate that explanations generally improve clinician trust and diagnostic confidence, but frequently increase cognitive load and exhibit misalignment with domain reasoning processes. To bridge these gaps, we propose a stakeholder-centric evaluation framework that integrates socio-technical principles and human-computer interaction to guide the future development of clinically viable and trustworthy XAI-based CDSS.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。