arXiv:2511.18732cs.LGstat.ML2025-11被引 4

构建首个开源全球海洋预报基准数据集,支持高效模型训练与评估。

OceanForecastBench: A Benchmark Dataset for Data-Driven Global Ocean Forecasting

  • 整合28年高精度再分析数据,覆盖4类变量23层深度
  • 基于超1亿个观测点实现可靠性能评估,涵盖海面与深海数据
  • 提供6个基线模型和完整评估流程,适合气候与海洋领域研究者

全球海洋预报旨在预测温度、盐度和洋流等关键海洋变量,对理解海洋现象至关重要。近年来,基于深度学习的数据驱动模型如XiHe、WenHai、LangYa和AI-GOMS在捕捉复杂海洋动力学和提升预报效率方面展现出显著潜力。然而,缺乏开源、标准化的基准数据集导致数据使用与评估方法不一致,阻碍了模型高效开发、公平比较及跨学科协作。为此,我们提出OceanForecastBench,包含三大核心贡献:(1) 28年全球海洋再分析数据,用于模型训练,包含4类海洋变量(23个深度层)及4类海表变量;(2) 高可靠性卫星与原位观测数据,覆盖约1亿个全球海洋位置,用于模型评估;(3) 评估流程与包含6个典型基线模型的综合基准框架,通过观测数据多维度评价模型表现。OceanForecastBench是当前最全面的数据驱动海洋预报基准平台,提供开源代码与数据,可支持模型开发、评估与对比。数据与代码已公开于:https://github.com/Ocean-Intelligent-Forecasting/OceanForecastBench。

原文摘要 · Abstract (English)

Global ocean forecasting aims to predict key ocean variables such as temperature, salinity, and currents, which is essential for understanding and describing oceanic phenomena. In recent years, data-driven deep learning-based ocean forecast models, such as XiHe, WenHai, LangYa and AI-GOMS, have demonstrated significant potential in capturing complex ocean dynamics and improving forecasting efficiency. Despite these advancements, the absence of open-source, standardized benchmarks has led to inconsistent data usage and evaluation methods. This gap hinders efficient model development, impedes fair performance comparison, and constrains interdisciplinary collaboration. To address this challenge, we propose OceanForecastBench, a benchmark offering three core contributions: (1) A high-quality global ocean reanalysis data over 28 years for model training, including 4 ocean variables across 23 depth levels and 4 sea surface variables. (2) A high-reliability satellite and in-situ observations for model evaluation, covering approximately 100 million locations in the global ocean. (3) An evaluation pipeline and a comprehensive benchmark with 6 typical baseline models, leveraging observations to evaluate model performance from multiple perspectives. OceanForecastBench represents the most comprehensive benchmarking framework currently available for data-driven ocean forecasting, offering an open-source platform for model development, evaluation, and comparison. The dataset and code are publicly available at: https://github.com/Ocean-Intelligent-Forecasting/OceanForecastBench.

海洋预报数据集基准测试深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。