用扩散模型从1%观测数据重建全球温度场,解决气候预测中的数据稀疏难题。
Probabilistic Spatial Interpolation of Sparse Data using Diffusion Models
- 基于预插值掩码引导的扩散模型,从极稀疏观测中恢复完整温度场。
- 在1%观测覆盖率下仍实现高精度重建,验证数据缺失场景下的鲁棒性。
- 适用于历史重分析与实时预报,特别适合观测稀疏区域的气候建模。
当前气候模型依赖于一个‘可信’的初始状态,即地球的合理快照,所有未来预测均由此出发。然而,由于系统内在的混沌特性,初始条件的微小不确定性会随时间呈指数级放大。这一挑战在大尺度和百年时间跨度下尤为突出,数据空缺不仅常见且不可避免。不确定性来源包括:(1) 卫星和地面站提供的稀疏、嘈杂观测;(2) 模型简化假设引发的内部变异性。实践中,数据同化方法通过将模型状态与部分观测条件结合来弥补信息缺失。本文工作在此基础上,聚焦极端稀疏场景,提出一种条件数据补全框架,仅需1%的观测覆盖率即可重建完整的温度场。该方法利用扩散模型并由预插值掩码引导,有效从极少数据点推断出全状态场。我们在美国南部大平原地区验证了该框架,聚焦2018–2020年夏季午后(12:00–18:00)的温度场,覆盖从条带数据到孤立站点的不同观测密度。结果表明,模型在各种稀疏条件下均表现出优异重建精度,展现了其在历史重分析与实时预报流程中填补关键数据空白的巨大潜力。
原文摘要 · Abstract (English)
The large underlying assumption of climate models today relies on the basis of a "confident" initial condition, a reasonably plausible snapshot of the Earth for which all future predictions depend on. However, given the inherently chaotic nature of our system, this assumption is complicated by sensitive dependence, where small uncertainties in initial conditions can lead to exponentially diverging outcomes over time. This challenge is particularly salient at global spatial scales and over centennial timescales, where data gaps are not just common but expected. The source of uncertainty is two-fold: (1) sparse, noisy observations from satellites and ground stations, and (2) internal variability stemming from the simplifying approximations within the models themselves. In practice, data assimilation methods are used to reconcile this missing information by conditioning model states on partial observations. Our work builds on this idea but operates at the extreme end of sparsity. We propose a conditional data imputation framework that reconstructs full temperature fields from as little as 1% observational coverage. The method leverages a diffusion model guided by a prekriged mask, effectively inferring the full-state fields from minimal data points. We validate our framework over the Southern Great Plains, focusing on afternoon (12:00-6:00 PM) temperature fields during the summer months of 2018-2020. Across varying observational densities--from swath data to isolated in-situ sensors--our model achieves strong reconstruction accuracy, highlighting its potential to fill in critical data gaps in both historical reanalysis and real-time forecasting pipelines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。