arXiv:2605.05854cs.AI2026-05被引 1

构建真实全球空气质量预测评估基准,揭示模型在碎片化数据下的真实表现。

AirQualityBench: A Realistic Evaluation Benchmark for Global Air Quality Forecasting

论文配图:AirQualityBench: A Realistic Evaluation Benchmark for Global Air Quality Forecasting
图 1 · 摘自论文原文
  • 基于3720个站点真实观测,保留原始缺失模式,不人为补全数据
  • 覆盖六种主要污染物,评估结果在物理浓度尺度上报告误差
  • 适合关注实际部署、遮蔽感知和可解释性模型的研究者使用

空气质量预测模型通常在区域性的、预处理过的标准化数据集上评估,其中缺失值被移除或人工填补。此类评估简化了对比,却掩盖了真实监测网络中的关键挑战:全球覆盖不均、结构化缺失、污染物尺度异质性及部署成本。我们提出 extbf{AirQualityBench},一个全球多污染物基准,用于在这些真实条件下评估预测模型。该基准包含2021—2025年期间3,720个监测站的小时级观测数据,覆盖六种主要污染物,并保留各数据提供方的原生观测掩码。不构造密集数据张量,而是将缺失性作为预测问题的一部分,误差在逆变换回物理浓度尺度后报告有效未来观测。在统一协议下评估代表性时空模型发现,强于清洗数据集的表现无法可靠迁移到全球碎片化监测流中。因此,AirQualityBench 成为可扩展、遮蔽感知且物理可解释的空气质量预测的现实测试平台。所有基准数据、代码、评估脚本及基线实现均可在 GitHub 获取。

原文摘要 · Abstract (English)

Air-quality forecasting models are commonly evaluated on regional, preprocessed, and normalized datasets, where missing observations are removed or artificially completed. Such protocols simplify comparison but hide the conditions that dominate real monitoring networks: uneven global coverage, structured missingness, heterogeneous pollutant scales, and deployment cost. We introduce \textbf{AirQualityBench}, a global multi-pollutant benchmark designed to evaluate forecasting models under these realistic conditions. The benchmark contains hourly observations from 3,720 monitoring stations over 2021--2025, covers six major pollutants, and preserves provider-native observation masks. Rather than imputing a dense data tensor, AirQualityBench exposes missingness as part of the forecasting problem and reports errors on valid future observations after inverse transformation to physical concentration scales. Evaluating representative spatio-temporal models under this unified protocol shows that strong performance on sanitized datasets does not reliably transfer to global, fragmented monitoring streams. AirQualityBench therefore serves as a realistic testbed for scalable, mask-aware, and physically interpretable air-quality forecasting. All benchmark data, code, evaluation scripts, and baseline implementations are available at \href{https://github.com/Star-Learning/AirQualityBench}{GitHub}.

空气质量真实评估缺失数据多污染物

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。