深度学习后处理方法在小样本时可能失效,新方法确保不同集合大小下预测公平可靠。
Ensemble-size-dependence of deep-learning post-processing methods that minimize an (un)fair score: motivating examples and a proof-of-concept solution
- 用变换器自注意力建模时间序列,保持成员独立性以满足公平评分要求。
- 传统方法因成员间耦合导致集合大小敏感,出现过度离散的系统偏差。
- 适用于气象预报等小集合场景,尤其适合对可靠性要求高的实际应用。
公平评分奖励与观测分布一致的集合成员,适合作为训练数据驱动集合预报或后处理方法的损失函数,尤其在大规模训练集合不可用或计算成本过高时。调整后的连续概率评分(aCRPS)在成员可交换且可视为从潜在预测分布中条件独立抽取时是公平且无偏的。然而,引入成员间结构依赖的分布感知后处理方法会破坏该假设,使aCRPS变得不公平。本文通过两种最小化有限集合期望aCRPS的方法验证此问题:(1) 基于集合均值耦合的线性逐成员校准;(2) 通过变换器跨集合维度自注意力耦合成员的深度学习方法。两者结果均对集合大小敏感,看似提升的aCRPS可能对应系统性不可靠性,表现为过度离散。为此提出轨迹变换器作为概念验证方案,其为PoET框架的改进版本,自注意力作用于预报时效而非集合维度,保持aCRPS所需的条件独立性。应用于欧洲中期天气预报中心(ECMWF)次季节预报系统的每周平均地表温度($T_{2m}$)预报,该方法在训练集为3或9成员、实时预报使用9或100成员时,均有效降低系统偏差并维持或提升预报可靠性。
原文摘要 · Abstract (English)
Fair scores reward ensemble forecast members that behave like samples from the same distribution as the verifying observations. They are therefore an attractive choice as loss functions to train data-driven ensemble forecasts or post-processing methods when large training ensembles are either unavailable or computationally prohibitive. The adjusted continuous ranked probability score (aCRPS) is fair and unbiased with respect to ensemble size, provided forecast members are exchangeable and interpretable as conditionally independent draws from an underlying predictive distribution. However, distribution-aware post-processing methods that introduce structural dependency between members can violate this assumption, rendering aCRPS unfair. We demonstrate this effect using two approaches designed to minimize the expected aCRPS of a finite ensemble: (1) a linear member-by-member calibration, which couples members through a common dependency on the sample ensemble mean, and (2) a deep-learning method, which couples members via transformer self-attention across the ensemble dimension. In both cases, the results are sensitive to ensemble size and apparent gains in aCRPS can correspond to systematic unreliability characterized by over-dispersion. We introduce trajectory transformers as a proof-of-concept that ensemble-size independence can be achieved. This approach is an adaptation of the Post-processing Ensembles with Transformers (PoET) framework and applies self-attention over lead time while preserving the conditional independence required by aCRPS. When applied to weekly mean $T_{2m}$ forecasts from the ECMWF subseasonal forecasting system, this approach successfully reduces systematic model biases whilst also improving or maintaining forecast reliability regardless of the ensemble size used in training (3 vs 9 members) or real-time forecasts (9 vs 100 members).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。