用机器学习快速生成美国各县慢性病估计,解决数据滞后问题
Geographically Weighted Surrogate Models for Rapid Small-Area Chronic Disease Estimation
- 基于地理加权机器学习,融合高频区域数据预测慢性病
- 在10种慢性病上表现优于全局模型,尤其在数据稀疏地区
- 适合公共卫生决策者用于实时健康态势监测
小区域估计(SAE)可帮助研究人员和政策制定者识别健康结果的空间差异,但基于调查的SAE产品存在固有延迟。以CDC PLACES为代表的权威估计通常在调查数据收集后约两年才发布,限制了其在时效性决策中的应用。本研究评估了机器学习(ML)作为替代方法的潜力:通过学习频繁更新的区域级预测因子与现有SAE输出之间的关系,在调查数据未发布或延迟的年份,生成及时且可比较的估计。我们评估了多种全局与地理加权机器学习模型,在全美各县对十种慢性病(包括慢阻肺、哮喘、心脏病、关节炎、癌症、抑郁、糖尿病、高血压、高胆固醇和中风)进行县一级的SAE。结果表明,地理加权随机森林和地理加权回归等框架能提供可扩展、开放数据支持的快速估计,有助于推动数据驱动的决策。
原文摘要 · Abstract (English)
Small-area estimation (SAE) enables researchers and policymakers to identify spatial disparities in health outcomes, but survey-based SAE products carry an inherent lag. Gold-standard estimates such as CDC PLACES are released roughly two years after the underlying survey data are collected, limiting their use for time-sensitive decision-making. This study evaluates the potential for machine learning (ML) to serve as a surrogate, learning the relationship between frequently updated area-level predictors and existing SAE outputs to generate timely, comparable estimates in years when SAE from surveys are unavailable or delayed. We evaluate several global and geographically weighted ML models for county-level SAE of ten chronic conditions across the US: COPD, asthma, heart disease, arthritis, cancer, depression, diabetes, high blood pressure, high cholesterol, and stroke. Our findings suggest that geographically weighted ML frameworks like geographically weighted random forest and geographically weighted regression offer scalable and open data surrogates for rapidly generating SAE and supporting data driven decision making.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。