用机器学习预测痴呆,准确率达98%。
Predictive Analytics for Dementia: Machine Learning on Healthcare Data
- 采用LDA等算法结合SMOTE处理数据不平衡问题。
- 模型测试准确率达98%,关键特征包括APOE-ε4基因和糖尿病。
- 强调可解释性AI对临床应用的重要性。
痴呆是一种影响认知与情绪功能的复杂综合征,阿尔茨海默病是最常见类型。本研究聚焦于利用机器学习技术对患者健康数据进行痴呆预测。采用监督学习算法,包括K近邻(KNN)、二次判别分析(QDA)、线性判别分析(LDA)和高斯过程分类器。为应对类别不平衡并提升模型性能,引入了合成少数类过采样技术(SMOTE)和词频-逆文档频率(TF-IDF)向量化方法。在所有模型中,LDA取得最高测试准确率98%。研究强调模型可解释性的重要性,并发现痴呆与载脂蛋白E-ε4等位基因存在关联,且慢性病如糖尿病也具显著相关性。本文呼吁未来机器学习研究应融合可解释人工智能方法,以进一步提升痴呆预测能力。
原文摘要 · Abstract (English)
Dementia is a complex syndrome impacting cognitive and emotional functions, with Alzheimer's disease being the most common form. This study focuses on enhancing dementia prediction using machine learning (ML) techniques on patient health data. Supervised learning algorithms are applied in this study, including K-Nearest Neighbors (KNN), Quadratic Discriminant Analysis (QDA), Linear Discriminant Analysis (LDA), and Gaussian Process Classifiers. To address class imbalance and improve model performance, techniques such as Synthetic Minority Over-sampling Technique (SMOTE) and Term Frequency-Inverse Document Frequency (TF-IDF) vectorization were employed. Among the models, LDA achieved the highest testing accuracy of 98%. This study highlights the importance of model interpretability and the correlation of dementia with features such as the presence of the APOE-epsilon4 allele and chronic conditions like diabetes. This research advocates for future ML innovations, particularly in integrating explainable AI approaches, to further improve predictive capabilities in dementia care.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。