对比多种模型,找到适合德州空气质量预测的最佳方案
Benchmarking Scientific Machine Learning Models for Air Quality Data
- 用物理规律约束深度学习模型,提升预测稳定性
- 短时预测中PM2.5和臭氧的精度提升最明显
- 为城市空气质量建模提供可复用的实证参考
准确预测空气质量指数(AQI)对快速发展的城市保护公众健康至关重要,但实际模型评估常受限于缺乏标准化、区域特定的基准测试。本研究构建了一个可解释的全面基准,对比经典时间序列、机器学习与深度学习方法在北德克萨斯州(达拉斯县)多时距AQI预测的表现。基于2022至2024年美国环保署(EPA)公开的每日空气质量数据,按城市聚合站点观测值,构建滞后量为{1,7,14,30}天的预测数据集,涵盖PM2.5和O3。评估了线性回归(LR)、SARIMAX、多层感知机(MLP)和LSTM网络,以及引入物理引导的变体(MLP+Physics和LSTM+Physics),后者通过加权损失将EPA断点法的AQI公式作为一致性约束。采用时间顺序训练-测试划分和MAE、RMSE误差指标,结果表明:深度学习模型优于简单基线,物理引导显著提升稳定性并确保污染物与AQI关系的物理一致性,尤其在短时预测及PM2.5和臭氧上效果最显著。整体为北德克萨斯地区模型选型提供实用依据,并明确轻量级物理约束在不同污染物与预测时距下的增益条件。
原文摘要 · Abstract (English)
Accurate air quality index (AQI) forecasting is essential for the protecting public health in rapidly growing urban regions, and the practical model evaluation and selection are often challenged by the lack of rigorous, region-specific benchmarking on standardized datasets. Physics-guided machine learning and deep learning models could be a good and effective solution to resolve such issues with more accurate and efficient AQI forecasting. This research study presents an explainable and comprehensive benchmark that enables a guideline and proposed physics-guided best model by benchmarking classical time-series, machine-learning, and deep-learning approaches for multi-horizon AQI forecasting in North Texas (Dallas County). Using publicly available U.S. Environmental Protection Agency (EPA) daily observations of air quality data from 2022 to 2024, we curate city-level time series for PM2.5 and O3 by aggregating station measurements and constructing lag-wise forecasting datasets for LAG in {1,7,14,30} days. For benchmarking the best model, linear regression (LR), SARIMAX, multilayer perceptrons (MLP), and LSTM networks are evaluated with the proposed physics-guided variants (MLP+Physics and LSTM+Physics) that incorporate the EPA breakpoint-based AQI formulation as a consistency constraint through a weighted loss. Experiments using chronological train-test splits and error metrics MAE, RMSE showed that deep-learning models outperform simpler baselines, while physics guidance improves stability and yields physically consistent pollutant with AQI relationships, with the largest benefits observed for short-horizon prediction and for PM2.5 and O3. Overall, the results provide a practical reference for selecting AQI forecasting models in North Texas and clarify when lightweight physics constraints meaningfully improve predictive performance across pollutants and forecast horizons.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。