系统梳理医疗AI的鲁棒性与可解释性,助力可信部署
Trustworthy AI in Digital Health: A Comprehensive Review of Robustness and Explainability

- 构建医疗AI可信框架,整合鲁棒性与可解释性方法
- 覆盖重症、新生儿等场景,强调真实世界应用挑战
- 适合研究者与临床工程师参考,尤其关注模型可靠性
确保AI系统的可信性对高风险领域如数字健康中的机器学习系统安全伦理集成至关重要。需在从问题定义、数据收集到模型部署与人机交互的全生命周期中,综合考虑鲁棒性、可解释性、公平性、问责制和隐私保护等维度。尽管已有诸多研究聚焦可信AI的特定方面,但针对医疗场景下鲁棒性与可解释性的系统性综述仍较匮乏。本文通过结构化梳理近年进展,提出一个面向数字健康的可信AI框架,涵盖技术与实践考量。第一部分介绍可信AI核心支柱及医疗领域的技术与伦理挑战;第二部分分析重症监护、新生儿健康、代谢疾病等具体应用场景中的信任需求;第三部分总结提升数据稀缺与分布偏移下鲁棒性的方法,以及基于特征重要性、梯度解释和反事实分析的可解释性技术。此外,还深入讨论了大模型时代可信AI的发展、评估指标(如有效性、保真度、多样性)及其在测量信任方面的应用。
原文摘要 · Abstract (English)
Ensuring trust in AI systems is essential for the safe and ethical integration of machine learning systems into high-stakes domains such as digital health. Key dimensions, including robustness, explainability, fairness, accountability, and privacy, need to be addressed throughout the AI lifecycle, from problem formulation and data collection to model deployment and human interaction. While various contributions address different aspects of trustworthy AI, a focused synthesis on robustness and explainability, especially tailored to the healthcare context, remains limited. This review addresses that need by organizing recent advancements into an accessible framework, highlighting both technical and practical considerations. We present a structured overview of methods, challenges, and solutions, aiming to support researchers and practitioners in developing reliable and explainable AI solutions for digital health. This review article is organized into three main parts. First, we introduce the pillars of trustworthy AI and discuss the technical and ethical challenges, particularly in the context of digital health. Second, we explore application-specific trust considerations across domains such as intensive care, neonatal health, and metabolic health, highlighting how robustness and explainability support trust. Lastly, we present recent advancements in techniques aimed at improving robustness under data scarcity and distributional shifts, as well as explainable AI methods ranging from feature attribution to gradient-based interpretations and counterfactual explanations. This paper is further enriched with detailed discussions of the contributions toward robustness and explainability in digital health, the development of trustworthy AI systems in the era of LLMs, and various evaluation metrics for measuring trust and related parameters such as validity, fidelity, and diversity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。