arXiv:2606.03995cs.LGcs.AI2026-06

用常规临床指标实现阿尔茨海默病三类分型,准确率超94%且可解释。

Early Detection of Alzheimer's Disease Using Explainable Machine Learning on Clinical Biomarkers: A Multi-Class Classification Study Using the Alzheimer's Disease Neuroimaging Initiative (ADNI) Dataset

论文配图:Early Detection of Alzheimer's Disease Using Explainable Machine Learning on Clinical Biomarkers: A Multi-Class Classification Study Using the Alzheimer's Disease Neuroimaging Initiative (ADNI) Dataset
图 1 · 摘自论文原文
  • 基于8个临床特征构建可解释的XGBoost分类器。
  • 在1641人数据上三分类准确率达94.4%,AUC达0.983。
  • 通过SHAP揭示不同阶段关键预测指标,助力临床决策。

阿尔茨海默病影响全球超5500万人。从常规临床评估中准确、可解释地区分正常认知(NC)、轻度认知障碍(MCI)和阿尔茨海默病(AD),仍是未满足的重大需求。本研究使用阿尔茨海默病神经影像学计划(ADNI)数据集中的8项临床指标(MMSE、CDR Global、CDR-SB、MoCA、FAQ、年龄、性别、教育程度),采用XGBoost分类器进行三分类。通过Optuna优化超参数(50次试验),并用SMOTE处理类别不平衡。性能评估采用五折交叉验证及1000次自助抽样95%置信区间,指标包括宏平均AUC-ROC、宏平均F1、平衡准确率与Cohen's kappa。SHAP值提供特征级可解释性。数据集共包含1641名基线受试者(NC 608例,MCI 767例,AD 266例)。五折交叉验证下,平均宏AUC为0.983(标准差0.007),准确率为0.944(标准差0.006),宏平均F1为0.929(标准差0.008)。在独立测试集(n=247)上,宏AUC为0.982(95% CI: 0.965–0.995),准确率0.943,平衡准确率0.932,宏平均F1 0.927,Cohen's kappa 0.909。SHAP分析表明,CDR Global是区分NC与MCI的主要预测因子,而CDR-SB与MMSE共同主导了AD分类。结论:基于常规临床评估的可解释机器学习模型实现了近乎完美的三类阿尔茨海默病检测。SHAP分析揭示了符合临床逻辑的类别特异性特征重要性模式,支持其临床有效性。未来工作将引入语音生物标志物,拓展多模态检测框架。

原文摘要 · Abstract (English)

Background: Alzheimer's disease (AD) affects over 55 million people worldwide. Accurate, interpretable detection of normal cognition (NC), mild cognitive impairment (MCI), and AD from routine clinical assessments remains a critical unmet need. Methods: An XGBoost classifier was developed for three-class detection using eight clinical features from the Alzheimer's Disease Neuroimaging Initiative (ADNI): MMSE, CDR Global, CDR Sum of Boxes (CDR-SB), MoCA, FAQ, age, sex, and education. Hyperparameters were optimised using Optuna (50 trials); class imbalance was addressed with SMOTE. Performance was evaluated by macro AUC-ROC with 1,000-iteration bootstrap 95% confidence intervals, macro F1, balanced accuracy, and Cohen's kappa. SHAP values provided feature-level explainability. Results: The dataset comprised 1,641 baseline subjects (608 NC, 767 MCI, 266 AD). On five-fold cross-validation, mean macro AUC was 0.983 (SD 0.007), accuracy 0.944 (SD 0.006), and macro F1 0.929 (SD 0.008). On the held-out test set (n = 247), macro AUC was 0.982 (95% CI: 0.965--0.995), accuracy 0.943, balanced accuracy 0.932, macro F1 0.927, and Cohen's kappa 0.909. SHAP analysis identified CDR Global as the dominant predictor for NC and MCI, while CDR-SB and MMSE together drove AD classification. Conclusion: An explainable machine learning model trained on routine clinical assessments achieves near-perfect three-class Alzheimer's detection. SHAP analysis reveals clinically plausible, class-specific feature importance patterns supporting clinical validity. Future work will extend this framework with speech biomarkers for multimodal detection.

阿尔茨海默病可解释AI临床诊断多分类

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。