用远程交通数据提升伦敦空气质量模型精度,尤其在交通密集区效果显著。
Leveraging Remote Traffic Data for Local Air Pollutant Estimation: A Scenario-Based Machine Learning Study Across London Monitoring Sites

- 用四种树模型结合交通、气象和时间变量预测污染物浓度。
- 加了交通数据后,NO₂预测误差降低至8.72–11.52 μg/m³。
- 交通变量贡献度接近邻站监测数据,适合城市污染建模研究者。
机动车排放是空气污染的主要来源,但远程获取的交通信息对本地机器学习空气质量模型的贡献尚未充分评估。本研究在伦敦多个监测点,对比四种可解释的树模型(随机森林、极深森林、LightGBM、XGBoost),在六种不同预测因子组合下预测NO₂、PM₁₀、PM₂.₅和O₃浓度,包括仅使用远程交通、气象与时间变量,以及加入1个或4个邻近站点数据的情况。以岭回归为基准,比较了空间插值方法和跨站点验证结果。当不使用邻近站点数据时,无交通信息下NO₂的均方根误差(RMSE)为9.73–11.66 μg/m³,加入交通信息后降至8.72–11.52 μg/m³。SHAP分析显示,在交通主导区域,交通相关变量的贡献可与邻近站点监测数据相媲美。
原文摘要 · Abstract (English)
Vehicular traffic is a major source of air pollution; however, the contribution of remotely acquired traffic information to local machine-learning (ML) air-pollution models remains insufficiently characterised. This study evaluates four interpretable tree-based ML models (Random Forest, Extra Trees, LightGBM, and XGBoost) under six predictor scenarios combining progressively larger predictor sets, ranging from remotely acquired traffic, meteorological, and temporal variables alone to the inclusion of measurements from one and four neighbouring monitoring stations, to estimate NO$_2$, PM$_{10}$, PM$_{2.5}$, and O$_3$ concentrations across several sites in London. ML model performance was compared with a ridge linear regression model as a baseline, with spatial interpolation methods and with a cross-site validation experiment. When modelling without data from neighbouring stations, the RMSE for NO$_2$ ranged from 9.73 to 11.66 $μ$g/m$^3$ without traffic information, compared with 8.72 to 11.52 $μ$g/m$^3$ when traffic information was included. Additionally, for NO$_2$, SHAP analyses indicate that traffic-related variables can contribute at levels comparable to pollutant measurements from neighbouring monitoring stations in traffic-dominated~environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。