用可解释AI发现多发性硬化症新生物标志物,比传统方法更全面。
A Machine Learning Pipeline for Multiple Sclerosis Biomarker Discovery: Comparing explainable AI and Traditional Statistical Approaches
- 融合8个公开基因数据集,用XGBoost与SHAP分析关键基因
- 发现30个独特于AI的潜在标志物,且与已知病理通路相关
- 适合想探索疾病机制的临床研究者和机器学习应用者
本研究构建了一套用于多发性硬化症(MS)生物标志物发现的机器学习流程,整合了来自外周血单个核细胞(PBMC)的八个公开微阵列数据集。经过严格预处理后,采用贝叶斯优化的XGBoost分类器进行训练,并利用SHapley Additive exPlanations(SHAP)识别对模型预测关键的特征,进而推断可能的生物标志物。这些结果与传统差异表达分析(DEA)识别的基因进行了比较,发现两者存在重叠和独立的标志物,表明方法互补。富集分析验证了SHAP筛选基因的生物学意义,其关联于鞘脂信号传导、Th1/Th2/Th17细胞分化以及爱泼斯坦-巴尔病毒(EBV)感染等已知与MS相关的通路。研究表明,结合可解释人工智能(xAI)与传统统计方法能更深入揭示疾病机制。
原文摘要 · Abstract (English)
We present a machine learning pipeline for biomarker discovery in Multiple Sclerosis (MS), integrating eight publicly available microarray datasets from Peripheral Blood Mononuclear Cells (PBMC). After robust preprocessing we trained an XGBoost classifier optimized via Bayesian search. SHapley Additive exPlanations (SHAP) were used to identify key features for model prediction, indicating thus possible biomarkers. These were compared with genes identified through classical Differential Expression Analysis (DEA). Our comparison revealed both overlapping and unique biomarkers between SHAP and DEA, suggesting complementary strengths. Enrichment analysis confirmed the biological relevance of SHAP-selected genes, linking them to pathways such as sphingolipid signaling, Th1/Th2/Th17 cell differentiation, and Epstein-Barr virus infection all known to be associated with MS. This study highlights the value of combining explainable AI (xAI) with traditional statistical methods to gain deeper insights into disease mechanism.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。