比较多种模型预测青少年肥胖,发现复杂算法提升有限。
Multilevel Determinants of Overweight and Obesity Among U.S. Children Aged 10-17: Comparative Evaluation of Statistical and Machine Learning Approaches Using the 2021 National Survey of Children's Health
- 用统计与机器学习模型分析1.8万多名青少年的肥胖影响因素。
- 复杂模型仅小幅提升召回率和精确度,整体效果不如简单逻辑回归稳定。
- 不同种族和贫困群体间预测差距持续存在,需关注数据公平性。
背景:美国儿童和青少年超重与肥胖仍是重大公共卫生问题,受行为、家庭和社区因素共同影响,其在人群层面的联合预测结构尚未完全明确。目标:识别美国10-17岁青少年超重/肥胖的多层级预测因子,并比较统计模型、机器学习与深度学习模型在预测性能、校准及亚组公平性方面的表现。数据与方法:基于2021年国家儿童健康调查数据,分析18,792名10-17岁儿童,采用体质指数(BMI)定义超重/肥胖。预测变量包括饮食、体力活动、睡眠、父母压力、社会经济状况、逆境经历及社区特征。评估模型涵盖逻辑回归、随机森林、梯度提升、XGBoost、LightGBM、多层感知机(MLP)和TabNet。性能指标包括AUC、准确率、精确率、召回率、F1分数和Brier评分。结果:模型区分能力范围为0.66至0.79。逻辑回归、梯度提升与MLP表现出最佳的区分与校准平衡。提升与深度学习模型略微改善召回率与F1分数。无任一模型全面优于其他。各族裔与贫困群体间的性能差异在各类算法中持续存在。结论:模型复杂度提升带来的增益有限,逻辑回归已具良好表现。关键预测因子始终覆盖行为、家庭与社区维度。亚组间持续存在的差距表明,亟需提升数据质量与开展以公平为导向的监测,而非追求更复杂的算法。
原文摘要 · Abstract (English)
Background: Childhood and adolescent overweight and obesity remain major public health concerns in the United States and are shaped by behavioral, household, and community factors. Their joint predictive structure at the population level remains incompletely characterized. Objectives: The study aims to identify multilevel predictors of overweight and obesity among U.S. adolescents and compare the predictive performance, calibration, and subgroup equity of statistical, machine-learning, and deep-learning models. Data and Methods: We analyze 18,792 children aged 10-17 years from the 2021 National Survey of Children's Health. Overweight/obesity is defined using BMI categories. Predictors included diet, physical activity, sleep, parental stress, socioeconomic conditions, adverse experiences, and neighborhood characteristics. Models include logistic regression, random forest, gradient boosting, XGBoost, LightGBM, multilayer perceptron, and TabNet. Performance is evaluated using AUC, accuracy, precision, recall, F1 score, and Brier score. Results: Discrimination range from 0.66 to 0.79. Logistic regression, gradient boosting, and MLP showed the most stable balance of discrimination and calibration. Boosting and deep learning modestly improve recall and F1 score. No model was uniformly superior. Performance disparities across race and poverty groups persist across algorithms. Conclusion: Increased model complexity yields limited gains over logistic regression. Predictors consistently span behavioral, household, and neighborhood domains. Persistent subgroup disparities indicate the need for improved data quality and equity-focused surveillance rather than greater algorithmic complexity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。