arXiv:2603.11725cs.CVcs.LG2026-03

用跨分辨率注意力提升高分辨率空气污染预测速度与精度

Cross-Resolution Attention Network for High-Resolution PM2.5 Prediction

  • 双分支视觉变换器融合低分辨率气象数据与高分辨率污染数据
  • 1公里分辨率欧洲地图预测仅需1.8秒,比基线减少4.7%~10.7%误差
  • 适合需要实时高精度环境监测的科研与公共政策应用

视觉变换器在时空预测中表现优异,但在超大尺度、高分辨率的环境监测场景中扩展性受限。以1公里分辨率绘制整个欧洲空气质量图包含2900万像素,远超普通自注意力计算极限。本文提出CRAN-PM,一种双分支视觉变换器,通过跨分辨率注意力机制高效融合25公里分辨率气象数据与当前时刻1公里分辨率PM2.5数据。不依赖温度、地形等物理输入,而是引入高程感知自注意力和风向引导的跨注意力,促使网络学习具有物理一致性的特征表示。该模型全可训练且内存高效,在单张GPU上1.8秒内完成2900万像素的欧洲地图生成。在2022年欧洲每日PM2.5预测任务(362天,2971个欧洲环境署监测站)中,相较最优单尺度基线,T+1时刻RMSE降低4.7%,T+3时刻降低10.7%,复杂地形区域偏差减少36%。

原文摘要 · Abstract (English)

Vision Transformers have achieved remarkable success in spatio-temporal prediction, but their scalability remains limited for ultra-high-resolution, continent-scale domains required in real-world environmental monitoring. A single European air-quality map at 1 km resolution comprises 29 million pixels, far beyond the limits of naive self-attention. We introduce CRAN-PM, a dual-branch Vision Transformer that leverages cross-resolution attention to efficiently fuse global meteorological data (25 km) with local high-resolution PM2.5 at the current time (1 km). Instead of including physically driven factors like temperature and topography as input, we further introduce elevation-aware self-attention and wind-guided cross-attention to force the network to learn physically consistent feature representations for PM2.5 forecasting. CRAN-PM is fully trainable and memory-efficient, generating the complete 29-million-pixel European map in 1.8 seconds on a single GPU. Evaluated on daily PM2.5 forecasting throughout Europe in 2022 (362 days, 2,971 European Environment Agency (EEA) stations), it reduces RMSE by 4.7% at T+1 and 10.7% at T+3 compared to the best single-scale baseline, while reducing bias in complex terrain by 36%.

PM2.5预测视觉变换器跨分辨率环境监测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。