融合个体与环境数据,用机器学习预测儿童肥胖风险
A Micro-Macro Machine Learning Framework for Predicting Childhood Obesity Risk Using NHANES and Environmental Determinants
- 构建微观-宏观机器学习框架,整合个人数据与环境指标
- XGBoost模型表现最佳,环境脆弱性指数与肥胖风险高度相关
- 适合公共卫生研究者用于风险识别与政策干预设计
儿童肥胖仍是美国重大公共卫生问题,受个体、家庭及环境等多层级因素影响。传统研究多孤立分析各层面,难以揭示结构性环境与个体特征的交互作用。本文提出一种微观-宏观机器学习框架,整合来自NHANES的个体体征与社会经济数据,以及美国农业部(USDA)和环保署(EPA)的宏观环境特征,包括食物可及性、空气质量与社会脆弱性。采用逻辑回归、随机森林、XGBoost和LightGBM四种模型进行肥胖预测,其中XGBoost表现最优。基于州级数据构建复合环境脆弱性指数(EnvScore),多层级对比显示高环境负担州与全国预测的微层级肥胖风险分布具有显著地理相似性。结果表明,跨尺度数据融合可有效识别环境驱动的肥胖风险差异。该工作提供可扩展的数据驱动建模流程,适用于公共卫生信息学,具备向因果建模、干预规划和实时分析拓展的潜力。
原文摘要 · Abstract (English)
Childhood obesity remains a major public health challenge in the United States, strongly influenced by a combination of individual-level, household-level, and environmental-level risk factors. Traditional epidemiological studies typically analyze these levels independently, limiting insights into how structural environmental conditions interact with individual-level characteristics to influence health outcomes. In this study, we introduce a micro-macro machine learning framework that integrates (1) individual-level anthropometric and socioeconomic data from NHANES and (2) macro-level structural environment features, including food access, air quality, and socioeconomic vulnerability extracted from USDA and EPA datasets. Four machine learning models Logistic Regression, Random Forest, XGBoost, and LightGBM were trained to predict obesity using NHANES microdata. XGBoost achieved the strongest performance. A composite environmental vulnerability index (EnvScore) was constructed using normalized indicators from USDA and EPA at the state level. Multi-level comparison revealed strong geographic similarity between states with high environmental burden and the nationally predicted micro-level obesity risk distribution. This demonstrates the feasibility of integrating multi-scale datasets to identify environment-driven disparities in obesity risk. This work contributes a scalable, data-driven, multi-level modeling pipeline suitable for public health informatics, demonstrating strong potential for expansion into causal modeling, intervention planning, and real-time analytics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。