解决气象数据降尺度中残差分布偏差问题,提升预测不确定性校准精度。
Mind the Residual Gap: Probabilistic Downscaling under Real-World Bias

- 通过低维主成分空间的最优传输对齐训练与测试时残差分布。
- 在合成数据和真实风场任务中显著降低欠分散,提升校准度(SSR、CRPS)。
- 适合需要高可靠性概率预测的气候建模与大气科学应用。
概率性降尺度旨在根据粗分辨率输入建模高分辨率场的条件分布,是大气科学、气候建模及其他多尺度物理系统的核心挑战。常用范式将问题分解为确定性均值预测器与随机残差生成器两部分。尽管在理想环境下有效,该方法在真实场景中常产生有偏且欠分散的集合预测。这是否仅为一般性预测不确定性校准偏差?我们发现根本原因更深层:残差目标误设——训练时诱导的残差分布与测试时所需分布因降尺度偏差而系统性不同。为此,我们提出 ReMatch(残差分布匹配),通过低维主成分分析空间中的最优传输,将训练残差分布向测试阶段对齐。该方法保留了均值-残差框架的统计优势,同时减少随机生成器在测试时面对的残差目标差异。在具有不同偏差水平的合成基准与真实世界 HRRR-ERA5 风场降尺度任务中,ReMatch 显著缓解欠分散,改善校准(SSR 和 CRPS),优于标准均值-残差模型及其变体,以及最先进的超分辨率模型。代码已公开于 https://github.com/sdean-group/ReMatch.git。
原文摘要 · Abstract (English)
Probabilistic downscaling is the task of modeling the conditional distribution of high-resolution fields given coarse inputs, and is a central challenge to atmospheric science, climate modeling, and other multiscale physical systems. A widely used paradigm decomposes the problem into a deterministic mean predictor followed by a stochastic residual generator. While effective in idealized settings, this mean--residual approach frequently produces biased and under-dispersive ensembles in real-world applications. Is this merely generic predictive uncertainty miscalibration? We show that the root cause is more fundamental: residual target misspecification, the residual distribution induced during training differs systematically from the one required at test time due to downscaling bias. To close this gap, we introduce ReMatch (Residual Distribution Matching). ReMatch aligns the training residual distribution toward the test-time regime via optimal transport in a low-dimensional PCA space. This preserves the statistical benefits of the mean--residual framework while reducing the train--test mismatch in the residual targets seen by the stochastic generator. On a controlled synthetic benchmark with varying bias levels and a real-world HRRR--ERA5 wind field downscaling task, ReMatch substantially reduces under-dispersion, improves calibration (SSR and CRPS), and outperforms strong baselines, including the standard mean--residual model and its variants, as well as state-of-the-art super-resolution models. Our code is available at https://github.com/sdean-group/ReMatch.git.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。