arXiv:2506.22447cs.LGcs.AI2025-06被引 1

用共享编码器+多解码器的ViT模型,同时预测六种气候变量,更准更快。

Vision Transformers for Multi-Variable Climate Downscaling: Emulating Regional Climate Models with a Shared Encoder and Multi-Decoder Architecture

  • 共享编码器+变量专用解码器,联合预测六类气候变量
  • 平均均方误差降低5.5%,推理速度比单变量快29-32%
  • 适合需要高精度与低算力的区域气候模拟场景

全球气候模型(GCM)虽能模拟大尺度气候动态,但空间分辨率粗,难以用于区域研究。区域气候模型(RCM)通过动力降尺度弥补此缺陷,但计算成本高且灵活性差。深度学习提供了高效的数据驱动替代方案,但现有方法多为单变量模型,逐变量处理导致计算冗余、上下文感知弱、变量间交互不足。为此,本文提出一种多变量视觉变换器(ViT)架构,采用共享编码器和变量专用解码器(1EMD),直接从GCM分辨率输入联合预测六项关键气候变量:地表温度、风速、500 hPa位势高度、总降水、地表向下短波辐射及地表向下长波辐射,实现对欧洲地区的高分辨率降尺度模拟。相较于单变量ViT模型,1EMD在所有六个变量上均表现更优,平均均方误差降低约5.5%。同时显著优于其他多变量基线模型,包括单解码器ViT和多变量U-Net。此外,多变量模型大幅降低计算开销,每变量推理时间减少29%-32%。结果表明,多变量建模在精度与效率上均有系统性优势,1EMD ViT在预测性能与计算成本间取得最优平衡。

原文摘要 · Abstract (English)

Global Climate Models (GCMs) are critical for simulating large-scale climate dynamics, but their coarse spatial resolution limits their applicability in regional studies. Regional Climate Models (RCMs) address this limitation through dynamical downscaling, albeit at considerable computational cost and with limited flexibility. Deep learning has emerged as an efficient data-driven alternative; however, most existing approaches focus on single-variable models that downscale one variable at a time. This paradigm can lead to redundant computation, limited contextual awareness, and weak cross-variable interactions.To address these limitations, we propose a multi-variable Vision Transformer (ViT) architecture with a shared encoder and variable-specific decoders (1EMD). The proposed model jointly predicts six key climate variables: surface temperature, wind speed, 500 hPa geopotential height, total precipitation, surface downwelling shortwave radiation, and surface downwelling longwave radiation, directly from GCM-resolution inputs, emulating RCM-scale downscaling over Europe. Compared to single-variable ViT models, the 1EMD architecture improves performance across all six variables, achieving an average MSE reduction of approximately 5.5% under a fair and controlled comparison. It also consistently outperforms alternative multi-variable baselines, including a single-decoder ViT and a multi-variable U-Net. Moreover, multi-variable models substantially reduce computational cost, yielding a 29-32% lower inference time per variable compared to single-variable approaches. Overall, our results demonstrate that multi-variable modeling provides systematic advantages for high-resolution climate downscaling in terms of both accuracy and efficiency. Among the evaluated architectures, the proposed 1EMD ViT achieves the most favorable trade-off between predictive performance and computational cost.

气候模拟Vision Transformer多变量建模降尺度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。