用街景图像和机器学习精准预测道路污染,关键在100米范围平均采样
How to predict on-road air pollution based on street view images and machine learning: a quantitative analysis of the optimal strategy
- 采集多角度街景图,100米缓冲区+平均策略提取特征
- 随机森林等机器学习模型误差低于2.5 μg/m³或ppb
- 解决过曝模糊问题,避免误判道路与人活动特征
道路空气污染在短距离内变化剧烈,受排放源、稀释及理化过程影响。融合移动监测数据与街景图像(SVIs)可提升局部污染预测能力。然而,算法选择、采样策略与图像质量引入额外误差,缺乏可靠参考量化其影响。为此,我们使用314辆出租车动态监测NO、NO2、PM2.5和PM10,并同步采集对应街景图像。从约38.2万张街景图中提取特征,覆盖0°、90°、180°、270°多个角度及100米至500米缓冲区。对比三种机器学习模型与线性土地利用回归(LUR)模型,结果表明:随机森林 > XGBoost > 神经网络 > LUR。相比单角度采样,平均策略能有效避免特征捕捉偏差。最优策略为100米缓冲区采样并采用平均特征提取,使各聚合点估计绝对误差几乎均小于2.5 μg/m³或ppb。过曝、模糊与欠曝导致图像误判,造成道路特征高估、人活动特征低估,进而影响污染物估算精度。
原文摘要 · Abstract (English)
On-road air pollution exhibits substantial variability over short distances due to emission sources, dilution, and physicochemical processes. Integrating mobile monitoring data with street view images (SVIs) holds promise for predicting local air pollution. However, algorithms, sampling strategies, and image quality introduce extra errors due to a lack of reliable references that quantify their effects. To bridge this gap, we employed 314 taxis to monitor NO, NO2, PM2.5 and PM10 dynamically and sampled corresponding SVIs, aiming to develop a reliable strategy. We extracted SVI features from ~ 382,000 streetscape images, which were collected at various angles (0°, 90°, 180°, 270°) and ranges (buffers with radii of 100m, 200m, 300m, 400m, 500m). Also, three machine learning algorithms alongside the linear land-used regression (LUR) model were experimented with to explore the influences of different algorithms. Four typical image quality issues were identified and discussed. Generally, machine learning methods outperform linear LUR for estimating the four pollutants, with the ranking: random forest > XGBoost > neural network > LUR. Compared to single-angle sampling, the averaging strategy is an effective method to avoid bias of insufficient feature capture. Therefore, the optimal sampling strategy is to obtain SVIs at a 100m radius buffer and extract features using the averaging strategy. This approach achieved estimation results for each aggregation location with absolute errors almost less than 2.5 μg/m^2 or ppb. Overexposure, blur, and underexposure led to image misjudgments and incorrect identifications, causing an overestimation of road features and underestimation of human-activity features, contributing to inaccurate NO, NO2, PM2.5 and PM10 estimation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。