arXiv:2508.07062physics.data-ancs.LG2025-08

规范气候预测中的数据预处理,提升模型可靠性

Setting the Standard: Recommended Practices for Data Preprocessing in Data-Driven Climate Prediction

  • 提出标准化异常、处理非平稳性等预处理标准方法
  • 不同预处理可导致相同模型输出差异,影响预测可信度
  • 适合气候建模者与AI应用开发者参考使用

人工智能(AI)和机器学习(ML)在跨时间尺度的气候预测中快速普及。由于数据驱动模型的性能直接受数据质量和预处理方式影响,因此亟需建立标准化的数据预处理规范。本文旨在:(1)帮助研究者理解预处理对气候预测的影响;(2)提出适用于次季节至十年以上气候预测的推荐预处理实践;(3)赋能最终用户判断所用模型是否适配其目标。内容涵盖标准化异常构建、非平稳性处理、时空相关性应对以及极端值与复杂分布变量的处理。案例表明,不同预处理策略会导致同一模型产生显著不同的预测结果,引发混淆并降低整体可信度。遵循本文建议可显著提升AI/ML在气候预测中的鲁棒性与透明度。

原文摘要 · Abstract (English)

Artificial intelligence (AI) - and specifically machine learning (ML) - applications for climate prediction across timescales are proliferating quickly. The emergence of these methods prompts a revisit to the impact of data preprocessing, a topic familiar to the climate community, as more traditional statistical models work with relatively small sample sizes. Indeed, the skill and confidence in the forecasts produced by data-driven models are directly influenced by the quality of the datasets and how they are treated during model development, thus yielding the colloquialism, "garbage in, garbage out." As such, this article establishes protocols for the proper preprocessing of input data for AI/ML models designed for climate prediction (i.e., subseasonal to decadal and longer). The three aims are to: (1) educate researchers, developers, and end users on the effects that preprocessing has on climate predictions; (2) provide recommended practices for data preprocessing for such applications; and (3) empower end users to decipher whether the models they are using are properly designed for their objectives. Specific topics covered in this article include the creation of (standardized) anomalies, dealing with non-stationarity and the spatiotemporally correlated nature of climate data, and handling of extreme values and variables with potentially complex distributions. Case studies will illustrate how using different preprocessing techniques can produce different predictions from the same model, which can create confusion and decrease confidence in the overall process. Ultimately, implementing the recommended practices set forth in this article will enhance the robustness and transparency of AI/ML in climate prediction studies.

气候预测数据预处理机器学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。