首个面向真实大气数据的多模态深度学习评估基准,解决模型比较不公问题。
DAMBench: A Multi-Modal Benchmark for Deep Learning-based Atmospheric Data Assimilation
- 构建真实气象背景与多源观测数据融合的统一评估平台
- 支持生成式模型与神经过程框架的公平对比,提升可复现性
- 适合气象建模、多模态深度学习研究者使用
数据同化是大气系统模拟的核心任务,旨在通过融合稀疏、含噪观测与先验估计来重构系统状态。传统方法如变分法和集合卡尔曼滤波虽有效,但近年来深度学习提供了更可扩展、高效且灵活的替代方案,尤其适用于包含大规模多模态观测的真实场景。然而现有深度学习同化研究存在两大局限:(1) 依赖合成扰动观测的简化场景;(2) 缺乏标准化基准进行公平模型对比。为此,本文提出DAMBench,首个大规模多模态数据同化基准,用于在真实大气条件下评估数据驱动模型。DAMBench整合了先进预报系统提供的高质量背景场及真实世界多模态观测(包括地面气象站与卫星图像),所有数据经重采样至统一网格并时间对齐,支持系统性训练、验证与测试。我们提供统一评估协议,并基准化代表性同化方法,包括潜在生成模型与神经过程框架。此外,提出轻量级多模态插件,展示如何通过引入真实观测提升简单基线性能。综合实验表明,DAMBench为未来研究奠定了严谨基础,推动可复现性、公平比较与真实场景扩展。数据与代码公开于 https://github.com/figerhaowang/DAMBench。
原文摘要 · Abstract (English)
Data Assimilation is a cornerstone of atmospheric system modeling, tasked with reconstructing system states by integrating sparse, noisy observations with prior estimation. While traditional approaches like variational and ensemble Kalman filtering have proven effective, recent advances in deep learning offer more scalable, efficient, and flexible alternatives better suited for complex, real-world data assimilation involving large-scale and multi-modal observations. However, existing deep learning-based DA research suffers from two critical limitations: (1) reliance on oversimplified scenarios with synthetically perturbed observations, and (2) the absence of standardized benchmarks for fair model comparison. To address these gaps, in this work, we introduce DAMBench, the first large-scale multi-modal benchmark designed to evaluate data-driven DA models under realistic atmospheric conditions. DAMBench integrates high-quality background states from state-of-the-art forecasting systems and real-world multi-modal observations (i.e., real-world weather stations and satellite imagery). All data are resampled to a common grid and temporally aligned to support systematic training, validation, and testing. We provide unified evaluation protocols and benchmark representative data assimilation approaches, including latent generative models and neural process frameworks. Additionally, we propose a lightweight multi-modal plugin to demonstrate how integrating realistic observations can enhance even simple baselines. Through comprehensive experiments, DAMBench establishes a rigorous foundation for future research, promoting reproducibility, fair comparison, and extensibility to real-world multi-modal scenarios. Our dataset and code are publicly available at https://github.com/figerhaowang/DAMBench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。