对比学习提升稀疏气象数据的低维表示,增强预测与极端天气检测性能。
Contrastive Learning Boosts Deterministic and Generative Models for Weather Data

- 通过对比损失对齐稀疏与完整数据,构建时空嵌入
- 在ERA5数据集上显著提升下游任务准确率,优于自编码器
- 融合物理知识图神经网络,适合真实稀疏气象场景
气象数据具有高维度和多模态特性,压缩为紧凑共享隐空间可提升下游任务效率。自监督学习中的对比学习能从无标签数据中生成鲁棒嵌入,尤其适用于标注数据稀缺的场景。尽管已有研究探索对比学习在ERA5数据上的应用,但缺乏与自编码器等压缩方法的系统比较,且未考虑实际采集中常见的稀疏数据问题。本文针对此提出SPARTA模型:通过对比损失对齐稀疏样本与完整样本,结合时序感知批采样策略与循环一致性损失优化隐空间结构,并引入新型图神经网络融合技术注入领域物理知识。实验表明,对比学习在稀疏气象数据下仍具优势,显著提升预测与极端天气检测性能。
原文摘要 · Abstract (English)
Weather data, comprising multiple variables, poses significant challenges due to its high dimensionality and multimodal nature. Creating low-dimensional embeddings requires compressing this data into a compact, shared latent space. This compression is required to improve the efficiency and performance of downstream tasks, such as forecasting or extreme-weather detection. Self-supervised learning, particularly contrastive learning, offers a way to generate low-dimensional, robust embeddings from unlabelled data, enabling downstream tasks when labelled data is scarce. Despite initial exploration of contrastive learning in weather data, particularly with the ERA5 dataset, the current literature does not extensively examine its benefits relative to alternative compression methods, notably autoencoders. Moreover, current work on contrastive learning does not investigate how these models can incorporate sparse data, which is more common in real-world data collection. It is critical to explore and understand how contrastive learning contributes to creating more robust embeddings for sparse weather data, thereby improving performance on downstream tasks. Our work extensively explores contrastive learning on the ERA5 dataset, aligning sparse samples with complete ones via a contrastive loss term to create SPARse-data augmented conTRAstive spatiotemporal embeddings (SPARTA). We introduce a temporally aware batch sampling strategy and a cycle-consistency loss to improve the structure of the latent space. Furthermore, we propose a novel graph neural network fusion technique to inject domain-specific physical knowledge. Ultimately, our results demonstrate that contrastive learning is a feasible and advantageous compression method for sparse geoscience data, thereby enhancing performance in downstream tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。