arXiv:2411.16421cs.CVcs.LG2024-11被引 7

扩充全球台风数据集,支持跨半球机器学习研究

Machine Learning for the Digital Typhoon Dataset: Extensions to Multiple Basins and New Developments in Representations and Tasks

  • 新增南半球台风数据,支持跨区域对比研究
  • 自监督学习+LSTM提升强度与转化预测性能
  • 新任务如中心定位,适合强台风检测场景

本文发布数字台风数据集V2,为40余年最长的台风卫星图像数据集,新增南半球热带气旋数据,与原北半球数据结合,支持跨半球区域差异研究。提出自监督学习框架用于表征学习,结合LSTM模型,在强度预报和温带转换预报任务中表现优异。引入台风中心估计等新任务,基于目标检测的模型在强台风上更优。通过在北半球训练、南半球测试,验证了模型跨半球泛化能力。数据集公开于 http://agora.ex.nii.ac.jp/digital-typhoon/dataset/ 及 https://github.com/kitamoto-lab/digital-typhoon/。

原文摘要 · Abstract (English)

This paper presents the Digital Typhoon Dataset V2, a new version of the longest typhoon satellite image dataset for 40+ years aimed at benchmarking machine learning models for long-term spatio-temporal data. The new addition in Dataset V2 is tropical cyclone data from the southern hemisphere, in addition to the northern hemisphere data in Dataset V1. Having data from two hemispheres allows us to ask new research questions about regional differences across basins and hemispheres. We also discuss new developments in representations and tasks of the dataset. We first introduce a self-supervised learning framework for representation learning. Combined with the LSTM model, we discuss performance on intensity forecasting and extra-tropical transition forecasting tasks. We then propose new tasks, such as the typhoon center estimation task. We show that an object detection-based model performs better for stronger typhoons. Finally, we study how machine learning models can generalize across basins and hemispheres, by training the model on the northern hemisphere data and testing it on the southern hemisphere data. The dataset is publicly available at \url{http://agora.ex.nii.ac.jp/digital-typhoon/dataset/} and \url{https://github.com/kitamoto-lab/digital-typhoon/}.

台风预测多半球自监督学习时序建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。