对比多种方法在82%缺失数据下预测空气质量,扩散模型表现最佳
Comparative Analysis of Machine Learning-Based Imputation Techniques for Air Quality Datasets with High Missing Data Rates
- 用外部特征增强扩散模型,提升高缺失率数据填补效果
- 扩散模型F1达0.9486,精度94.26%,在82.42%缺失率下表现最优
- 适合处理高缺失率的环境监测数据,尤其城市空气污染研究
城市污染对健康构成严重威胁,尤以交通相关空气污染为甚,对行人和骑行者等暴露人群影响显著。因此,高空间分辨率的空气质量监测对城市环境管理至关重要。本研究聚焦于处理缺失率高达82.42%的时空数据集,该数据来自都柏林的移动传感器与固定站点,由动态地块分发、环保局及谷歌联合采集。针对高缺失率导致的PM2.5精准分类困难,评估并比较了集成方法、深度学习模型与扩散模型等多种插补与预测技术。引入交通流量、气象条件及最近站点数据作为外部特征以提升性能。结果表明,结合外部特征的扩散模型取得最高F1分数(0.9486),准确率94.26%,精确率94.42%,召回率94.82%;而集成模型达到最高准确率94.82%,证明在极高缺失率下仍可实现优异性能。
原文摘要 · Abstract (English)
Urban pollution poses serious health risks, particularly in relation to traffic-related air pollution, which remains a major concern in many cities. Vehicle emissions contribute to respiratory and cardiovascular issues, especially for vulnerable and exposed road users like pedestrians and cyclists. Therefore, accurate air quality monitoring with high spatial resolution is vital for good urban environmental management. This study aims to provide insights for processing spatiotemporal datasets with high missing data rates. In this study, the challenge of high missing data rates is a result of the limited data available and the fine granularity required for precise classification of PM2.5 levels. The data used for analysis and imputation were collected from both mobile sensors and fixed stations by Dynamic Parcel Distribution, the Environmental Protection Agency, and Google in Dublin, Ireland, where the missing data rate was approximately 82.42%, making accurate Particulate Matter 2.5 level predictions particularly difficult. Various imputation and prediction approaches were evaluated and compared, including ensemble methods, deep learning models, and diffusion models. External features such as traffic flow, weather conditions, and data from the nearest stations were incorporated to enhance model performance. The results indicate that diffusion methods with external features achieved the highest F1 score, reaching 0.9486 (Accuracy: 94.26%, Precision: 94.42%, Recall: 94.82%), with ensemble models achieving the highest accuracy of 94.82%, illustrating that good performance can be obtained despite a high missing data rate.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。