arXiv:2410.15218cs.LG2024-10被引 3

深度学习模型融合外部数据可显著提升水文时间序列预测精度。

Deep Learning Foundation and Pattern Models: Challenges in Hydrological Time Series

  • 通过分析全球水文数据,探索多源时序特征的建模方法。
  • 引入外部变量使均方误差降低最高达40%。
  • 适合关注复杂时序建模的气象、水文研究人员。

深度学习在时间序列分析中备受关注,但多数研究未深入科学应用。本文聚焦水文时间序列,基于CAMELS与Caravan全球数据集,涵盖约8,000个流域、最多六条观测流和209个静态参数,分析降雨与径流数据。研究评估了八种不同模型配置下外部变量的影响,结果表明融合外部信息能显著提升表征能力,最大可使均方误差降低40%。对超过20种先进模式与基础模型进行性能对比,发现整合全面观测与外部数据的模型优于仅依赖有限输入或基础模型的方法。尤其自然年周期性外部时序贡献最显著,静态及其它周期因素亦具价值。所有分析代码开源,基于Google Colab的Jupyter Notebook实现,支持LSTM建模、数据预处理与模型比较。

原文摘要 · Abstract (English)

There has been active investigation into deep learning approaches for time series analysis, including foundation models. However, most studies do not address significant scientific applications. This paper aims to identify key features in time series by examining hydrology data. Our work advances computer science by emphasizing critical application features and contributes to hydrology and other scientific fields by identifying modeling approaches that effectively capture these features. Scientific time series data are inherently complex, involving observations from multiple locations, each with various time-dependent data streams and exogenous factors that may be static or time-varying and either application-dependent or purely mathematical. This research analyzes hydrology time series from the CAMELS and Caravan global datasets, which encompass rainfall and runoff data across catchments, featuring up to six observed streams and 209 static parameters across approximately 8,000 locations. Our investigation assesses the impact of exogenous data through eight different model configurations for key hydrology tasks. Results demonstrate that integrating exogenous information enhances data representation, reducing mean squared error by up to 40% in the largest dataset. Additionally, we present a detailed performance comparison of over 20 state-of-the-art pattern and foundation models. The analysis is fully open-source, facilitated by Jupyter Notebook on Google Colab for LSTM-based modeling, data preprocessing, and model comparisons. Preliminary findings using alternative deep learning architectures reveal that models incorporating comprehensive observed and exogenous data outperform more limited approaches, including foundation models. Notably, natural annual periodic exogenous time series contribute the most significant improvements, though static and other periodic factors are also valuable.

水文预测时间序列深度学习外部变量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。