arXiv:2506.13904cs.HCcs.AI2025-06综述被引 3

梳理医疗AI可解释性评估方法,给出用户中心设计指南。

A Systematic Review of User-Centred Evaluation of Explainable AI in Healthcare

  • 构建医疗XAI用户体验的原子化评价属性框架
  • 分析82项研究,发现人因评估正成主流趋势
  • 提供可落地的评估策略指南,适合跨学科团队使用

尽管可解释人工智能(XAI)取得进展,其实用价值在真实场景中仍缺乏充分探索与验证。可靠的、情境感知的评估至关重要,不仅需生成易懂的解释,还需确保其对目标用户的可信度与可用性,但因缺乏用户评估设计的明确指引,常被忽视。本研究旨在填补这一空白,有两个核心目标:(1) 构建一套定义清晰、原子化的属性框架,用于刻画医疗XAI中的用户体验;(2) 提供基于系统特性的、情境敏感的评估策略制定指南。我们对来自五个数据库的82项用户研究进行了系统综述,所有研究均聚焦于医疗场景中对AI生成解释的评估。分析基于预设编码方案,参考现有评估框架,并通过迭代发展归纳编码。研究得出三项关键贡献:(1) 汇总当前评估实践,揭示医疗XAI领域日益重视人因导向方法的趋势;(2) 揭示解释属性间的内在关联;(3) 更新了评估框架,并提出一系列可操作指南,支持跨学科团队为特定应用情境设计并实施有效的XAI评估策略。

原文摘要 · Abstract (English)

Despite promising developments in Explainable Artificial Intelligence, the practical value of XAI methods remains under-explored and insufficiently validated in real-world settings. Robust and context-aware evaluation is essential, not only to produce understandable explanations but also to ensure their trustworthiness and usability for intended users, but tends to be overlooked because of no clear guidelines on how to design an evaluation with users. This study addresses this gap with two main goals: (1) to develop a framework of well-defined, atomic properties that characterise the user experience of XAI in healthcare; and (2) to provide clear, context-sensitive guidelines for defining evaluation strategies based on system characteristics. We conducted a systematic review of 82 user studies, sourced from five databases, all situated within healthcare settings and focused on evaluating AI-generated explanations. The analysis was guided by a predefined coding scheme informed by an existing evaluation framework, complemented by inductive codes developed iteratively. The review yields three key contributions: (1) a synthesis of current evaluation practices, highlighting a growing focus on human-centred approaches in healthcare XAI; (2) insights into the interrelations among explanation properties; and (3) an updated framework and a set of actionable guidelines to support interdisciplinary teams in designing and implementing effective evaluation strategies for XAI systems tailored to specific application contexts.

可解释AI医疗AI用户评估人因设计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。