arXiv:2602.06800cs.LG2026-02

用流匹配技术实现低延迟高精度天气数据融合,提升预报效率与稳定性。

FlowDA: Accurate, Low-Latency Weather Data Assimilation via Flow Matching

  • 基于流匹配的生成式框架,通过集合卷积嵌入观测数据
  • 在0.1%~3.9%观测率下优于主流基线,参数量相当
  • 抗噪声强,长周期滚动融合稳定,适合实际业务应用

数据同化(DA)是现代天气预报的核心环节,但传统变分方法在机器学习预报流程中仍存在重大计算瓶颈。近期生成式机器学习方法虽具潜力,却普遍需要大量采样步骤,且在长时序自回归循环同化中易积累误差。本文提出FlowDA,一种基于流匹配的低延迟气象尺度生成式同化框架。FlowDA通过SetConv结构对观测进行嵌入,并微调Aurora基础模型,实现高效、准确、鲁棒的分析结果。在观测率从3.9%降至0.1%的多场景实验中,FlowDA在参数量相近的前提下显著优于多个强基线。同时,其对观测噪声具有强鲁棒性,在长时序自回归循环同化中表现稳定。总体表明,FlowDA为数据驱动同化提供了高效可扩展的新方向。

原文摘要 · Abstract (English)

Data assimilation (DA) is a fundamental component of modern weather prediction, yet it remains a major computational bottleneck in machine learning (ML)-based forecasting pipelines due to reliance on traditional variational methods. Recent generative ML-based DA methods offer a promising alternative but typically require many sampling steps and suffer from error accumulation under long-horizon auto-regressive rollouts with cycling assimilation. We propose FlowDA, a low-latency weather-scale generative DA framework based on flow matching. FlowDA conditions on observations through a SetConv-based embedding and fine-tunes the Aurora foundation model to deliver accurate, efficient, and robust analyses. Experiments across observation rates decreasing from $3.9\%$ to $0.1\%$ demonstrate superior performance of FlowDA over strong baselines with similar tunable-parameter size. FlowDA further shows robustness to observational noise and stable performance in long-horizon auto-regressive cycling DA. Overall, FlowDA points to an efficient and scalable direction for data-driven DA.

天气预测数据同化流匹配生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。