arXiv:2605.24945cs.LGcs.AI2026-05

RealBench让天气预报模型在真实操作条件下接受更严格的考验。

RealBench: Benchmarking Data-Driven Numerical Weather Forecasting Under Operational Conditions and Extreme Event Challenges

论文配图:RealBench: Benchmarking Data-Driven Numerical Weather Forecasting Under Operational Conditions and Extreme Event Challenges
图 1 · 摘自论文原文
  • 用2025年数据构建无泄露测试集,模拟真实预报环境。
  • 结合超1万站点观测数据,评估结果更贴近实际大气状态。
  • 专门评测极端天气,揭示传统榜单的严重偏差。

准确评估天气预报模型对其实用部署至关重要。然而,现有基准大多依赖如ERA5等再分析产品,这些数据通过延迟同化生成,无法反映实时预报的约束条件,导致基准性能与真实场景存在系统性偏差。本文提出RealBench,下一代面向人工智能天气预报的基准体系,强调在真实操作条件下的评估。该基准采用严格跨分布测试集,覆盖2025年,避免数据泄露并捕捉近期大气状态。集成低延迟业务分析数据与大规模全球地面观测数据集(超10,000个站点),可直接与真实大气测量值对比。除标准全球指标外,还提供针对热浪、寒潮、热带气旋等高影响极端事件的专项评估框架,使用事件特异性指标更贴合实际预报需求。评估结果显示,基于再分析数据的指标与真实表现存在显著差异,尤其在极端事件上。本工作揭示了现有基准的局限性,建立更忠实、更具操作意义的评估范式,为下一代AI天气预报系统发展提供严谨基础。代码已开源:https://github.com/lixruize-del/NWP-Benchmark。

原文摘要 · Abstract (English)

Accurate evaluation of weather forecasting models is critical for their reliable deployment in real-world applications. However, existing benchmarks predominantly rely on reanalysis products such as ERA5, which are generated through delayed data assimilation and do not reflect the constraints of real-time operational forecasting, thereby resulting in a systematic mismatch between benchmark performance and real-world forecasting. In this work, we introduce RealBench, a next-generation benchmark for AI weather forecasting that emphasizes realistic evaluation under operational conditions. RealBench features a strictly out-of-distribution test set spanning 2025 to eliminate data leakage and capture recent atmospheric regimes. It integrates multiple data sources, including low-latency operational analysis and a large-scale global in-situ observation dataset comprising over 10,000 stations, enabling direct evaluation against real atmospheric measurements. Beyond standard global metrics, RealBench provides a comprehensive evaluation framework for high-impact extreme events, including heatwaves, cold surges, and tropical cyclones, using event-specific metrics that better reflect real-world forecasting priorities. The evaluation results reveal substantial discrepancies between reanalysis-based metrics and real-world performance, particularly concerning extreme events. By highlighting the limitations of existing benchmarks, this work establishes a more faithful and operationally relevant evaluation paradigm, providing a rigorous foundation for advancing next-generation AI weather forecasting systems. The benchmark implementation is available at: https://github.com/lixruize-del/NWP-Benchmark.

天气预报基准测试极端事件

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。