用多维度指标验证阿尔茨海默病模型解释的稳定性,提升临床可信度。
Enhancing SHAP Explainability for Diagnostic and Prognostic ML Models in Alzheimer Disease
- 构建多层级可解释性框架,评估特征重要性与SHAP的一致性
- 认知和功能指标在诊断与预后中均主导解释结果,基因贡献微弱
- 跨任务、跨阶段解释高度稳定,适合临床决策支持
阿尔茨海默病(AD)的诊断与预后日益依赖机器学习(ML)模型,但其临床应用受限于技术门槛及缺乏可信、一致的模型解释。现有研究多关注单一任务的SHAP解释,缺乏对疾病阶段、模型架构或预测目标下解释鲁棒性的验证。本文提出一种多层级可解释性框架,通过整合:(1) 模型内特征重要性与SHAP的一致性度量;(2) 跨AD阶段的SHAP稳定性;(3) 诊断与预后任务间的跨任务一致性。基于NACC数据集,使用AutoML训练四类诊断与四类预后模型,覆盖标准AD进展阶段。通过相关性、top-k特征重叠、SHAP符号一致性及领域贡献率等指标评估稳定性。结果显示,认知与功能指标在诊断与预后中始终主导解释,遗传因素仅在预后中略有增加。所有分类器中,诊断与预后模型的SHAP-SHAP符号一致性达100%,解释量级变化极小。表明SHAP解释具备可量化验证的鲁棒性与可迁移性,为临床提供更可靠的预测解释。
原文摘要 · Abstract (English)
Alzheimer disease (AD) diagnosis and prognosis increasingly rely on machine learning (ML) models. Although these models provide good results, clinical adoption is limited by the need for technical expertise and the lack of trustworthy and consistent model explanations. SHAP (SHapley Additive exPlanations) is com-monly used to interpret AD models, but existing studies tend to focus on explanations for isolated tasks, providing little evidence about their robustness across disease stages, model architectures, or prediction objectives. This paper proposes a multi-level explainability framework that measures the coherence, stabil-ity and consistency of explanations by integrating: (1) within-model coherence metrics between feature importance and SHAP, (2) SHAP stability across AD boundaries, and (3) SHAP cross-task consistency be-tween diagnosis and prognosis. Using AutoML to optimize classifiers on the NACC dataset, we trained four diagnostic and four prognostic models covering the standard AD progression stages. Stability was then evaluated using correlation metrics, top-k feature overlap, SHAP sign consistency, and domain-level contribution ratios. Results show that cognitive and functional markers dominate SHAP explanations in both diagnosis and prognosis. SHAP-SHAP consistency between diagnostic and prognostic models was high across all classifiers, with 100% sign stability and minimal shifts in explanatory magnitude. Domain-level contributions also remained stable, with only minimal increases in genetic features for prognosis. These results demonstrate that SHAP explanations can be quantitatively vali-dated for robustness and transferability, providing clinicians with more reliable interpretations of ML pre-dictions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。