用机器学习预测尼泊尔5岁以下儿童贫血,关键特征可解释且效果佳。
Predicting Anemia Among Under-Five Children in Nepal Using Machine Learning and Deep Learning
- 融合四种特征选择法,筛选出5个核心风险因素。
- 逻辑回归召回率0.701,支持向量机AUC达0.736。
- 结果对基层筛查与政策制定具实用价值。
尼泊尔5岁以下儿童贫血仍是重大公共卫生问题,与生长发育迟缓、认知障碍及发病率升高相关。基于世卫组织血红蛋白阈值,将6-59月龄儿童的贫血状态定义为二分类任务(贫血 vs 非贫血)。分析2022年尼泊尔人口健康调查(NDHS 2022)微数据,涵盖1,855名儿童,初始考虑48项候选特征,包括人口统计、社会经济、母亲及儿童健康指标。采用卡方检验、互信息、点二列相关和Boruta四种特征选择方法,以多方法共识确定稳定特征集。最终一致选出5个特征:儿童年龄、近期发热、家庭规模、母亲贫血史及驱虫史;月经停止、族裔指标和省份也常被保留。随后比较8种传统机器学习模型(逻辑回归、KNN、决策树、随机森林、XGBoost、SVM、朴素贝叶斯、线性判别分析)与2种深度学习模型(DNN、TabNet),使用标准评估指标,重点考量F1-score与召回率以应对类别不平衡。结果显示,逻辑回归在召回率(0.701)和F1-score(0.649)上表现最佳,DNN准确率达0.709,而SVM的AUC最高(0.736)。总体表明,机器学习与深度学习模型均可实现具有竞争力的贫血预测能力,且儿童年龄、感染指征、母亲贫血与驱虫史等可解释特征在风险分层与公共卫生筛查中具有核心作用。
原文摘要 · Abstract (English)
Childhood anemia remains a major public health challenge in Nepal and is associated with impaired growth, cognition, and increased morbidity. Using World Health Organization hemoglobin thresholds, we defined anemia status for children aged 6-59 months and formulated a binary classification task by grouping all anemia severities as \emph{anemic} versus \emph{not anemic}. We analyzed Nepal Demographic and Health Survey (NDHS 2022) microdata comprising 1,855 children and initially considered 48 candidate features spanning demographic, socioeconomic, maternal, and child health characteristics. To obtain a stable and substantiated feature set, we applied four features selection techniques (Chi-square, mutual information, point-biserial correlation, and Boruta) and prioritized features supported by multi-method consensus. Five features: child age, recent fever, household size, maternal anemia, and parasite deworming were consistently selected by all methods, while amenorrhea, ethnicity indicators, and provinces were frequently retained. We then compared eight traditional machine learning classifiers (LR, KNN, DT, RF, XGBoost, SVM, NB, LDA) with two deep learning models (DNN and TabNet) using standard evaluation metrics, emphasizing F1-score and recall due to class imbalance. Among all models, logistic regression attained the best recall (0.701) and the highest F1-score (0.649), while DNN achieved the highest accuracy (0.709), and SVM yielded the strongest discrimination with the highest AUC (0.736). Overall, the results indicate that both machine learning and deep learning models can provide competitive anemia prediction and the interpretable features such as child age, infection proxy, maternal anemia, and deworming history are central for risk stratification and public health screening in Nepal.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。