用可解释模型分析污染与呼吸健康关系,发现低收入国家更依赖PM2.5预测。
Interpretable Machine Learning for Air Pollution and Respiratory Health Prediction: A Socioeconomic Subgroup Analysis
- 构建可解释机器学习框架,融合经济与气象数据预测呼吸疾病和空气质量。
- PM2.5是呼吸病率的主导因子,移除后分类准确率显著下降。
- 揭示低收入国家对PM2.5更敏感,凸显环境不平等影响健康预测。
空气污染与气候压力日益威胁呼吸健康,尤其在环境暴露与医疗资源不均的地区。本研究评估了一种可解释机器学习框架,基于国家层面的周度结构化数据,预测呼吸系统疾病发病率(每10万人口)和空气质量状态。采用两阶段监督学习:回归任务预测疾病率,二分类任务判断空气质量。对比了九种回归模型与九种分类模型,使用嵌套交叉验证评估性能。通过SHAP值进行模型解释,并按收入水平与地理区域开展子群分析。结果显示,PM2.5浓度是呼吸病率的主导预测因子,线性及正则化线性模型表现最优。空气质量分类中,包含PM2.5时模型达到高平衡准确率,但移除后性能大幅下降,表明高度依赖污染物信息。SHAP分析显示,去除PM2.5后,人均GDP、降水与医疗可及性等社会经济与气象变量影响增强。子群分析表明各收入组整体回归误差相似,但PM2.5对低收入中等收入国家的预测贡献更强。结果表明,仅关注模型精度不足以支撑气候-健康预测;可解释模型有助于识别关键污染信号,检验结果对核心污染物的依赖性,并揭示不同社会经济群体间的预测差异。
原文摘要 · Abstract (English)
Air pollution and climate-related stressors are increasingly important concerns for respiratory health, especially in settings with unequal environmental exposure and healthcare capacity. This study evaluates an interpretable machine learning framework for predicting respiratory disease rates and air-quality status using structured country-level weekly data. Two supervised learning tasks were considered: regression of respiratory disease rate per 100,000 population and binary classification of air-quality status. Nine regression models and nine classification models were compared using nested cross-validation. Model interpretation was conducted using SHAP values, and subgroup analysis was performed across income levels and geographic regions. The results showed that PM2.5 concentration was the dominant predictor of respiratory disease rate, with linear and regularized linear models achieving the strongest regression performance. For air-quality classification, models achieved high balanced accuracy when PM2.5 was included, but performance decreased substantially when PM2.5 was removed, indicating strong dependence on pollutant-related information. SHAP analysis showed that, without PM2.5, socioeconomic and meteorological variables such as GDP per capita, precipitation, and healthcare access became more influential. Subgroup analysis showed similar aggregate regression error across income groups, but PM2.5 contributed more strongly to predictions in lower-middle-income countries. These results show that model accuracy alone is not sufficient for climate-health prediction. Interpretable models can help identify dominant pollution-related signals, test whether results depend on key pollutant variables, and show whether prediction patterns differ across socioeconomic groups.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。