arXiv:2502.04409cs.LGphysics.ao-ph2025-02被引 3

用自编码器降低天气预报集合数据维度,保留概率特性。

Learning low-dimensional representations of ensemble forecast fields using autoencoder-based methods

  • 分成员处理后合并为概率分布,或联合编码生成低维表示
  • 在欧洲十年温风数据上验证,保持空间与统计特征
  • 适合需要压缩概率型气象数据的研究者

大规模数值模拟常产生高维网格数据,难以用于下游任务。以数值天气预报为例,大气过程通过离散网格化物理变量和动力学建模,不确定性通过多次随机模拟形成集合场,构成高维随机预测分布。集合数据的高维度和大体量给后续预测带来重大计算挑战。数据驱动的降维技术可学习有意义且紧凑的表示,减少数据量。然而,现有降维方法多针对确定性单值输入,无法处理多随机模拟产生的集合数据。本研究提出两种专为集合预报场设计的新降维方法:第一种对每个成员分别应用标准降维,再将结果合并为联合参数化分布模型;第二种通过定制变分自编码器联合编码所有成员。我们在欧洲十年温度与风速预报数据上进行案例研究,评估并比较两种方法。结果表明,两者均能保留集合的关键空间与统计特征,并支持概率性预报场重建。

原文摘要 · Abstract (English)

Large-scale numerical simulations often produce high-dimensional gridded data that is challenging to process for downstream applications. A prime example is numerical weather prediction, where atmospheric processes are modeled using discrete gridded representations of the physical variables and dynamics. Uncertainties are assessed by running the simulations multiple times, yielding ensembles of simulated fields as a high-dimensional stochastic representation of the forecast distribution. The high-dimensionality and large volume of ensemble datasets poses major computing challenges for subsequent forecasting stages. Data-driven dimensionality reduction techniques could help to reduce the data volume before further processing by learning meaningful and compact representations. However, existing dimensionality reduction methods are typically designed for deterministic and single-valued inputs, and thus cannot handle ensemble data from multiple randomized simulations. In this study, we propose novel dimensionality reduction approaches specifically tailored to the format of ensemble forecast fields. We present two alternative frameworks, which yield low-dimensional representations of ensemble forecasts while respecting their probabilistic character. The first approach derives a distribution-based representation of an input ensemble by applying standard dimensionality reduction techniques in a member-by-member fashion and merging the member representations into a joint parametric distribution model. The second approach achieves a similar representation by encoding all members jointly using a tailored variational autoencoder. We evaluate and compare both approaches in a case study using 10 years of temperature and wind speed forecasts over Europe. The approaches preserve key spatial and statistical characteristics of the ensemble and enable probabilistic reconstructions of the forecast fields.

降维集合预报自编码器气象建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。