用分散时间点的样本训练气候模型,效果优于连续历史数据。
Temporal Coverage over Density: Parsimonious Training-Set Design for ML Climate Downscaling

- 从整个气候变化轨迹中分散选取训练年份,而非只用历史连续年份。
- 仅用十分之一数据训练,仍可媲美全量数据模型性能。
- 适合需要高效利用计算资源的气候模拟与预测研究者。
高分辨率区域气候模拟对气候影响评估至关重要,但计算成本高昂,促使机器学习降尺度模型和代理模型的发展。关键挑战在于如何在有限计算预算下,将有限的高分辨率模拟数据分布在变化的气候轨迹中,以同时捕捉强迫响应和内部变率。基于西美地区CESM2大型集合的实验表明,在固定数据预算下,三种训练年份选择策略对比:连续历史年份、起始与末期年份混合、以及贯穿整个气候轨迹的分散年份。包含历史与未来年份的训练始终优于仅使用历史年份,说明暴露于历史记录之外的气候状态至关重要,揭示了统计降尺度中常见平稳性假设的局限性。贯穿全轨迹的分散采样表现最佳,表明广泛覆盖内部变率能提供额外信息,仅靠强迫响应无法满足需求。在未见集合成员中,分散采样训练模型更准确再现变率,且在多种气候诊断指标上保持优异性能。即使仅使用十分之一可用高分辨率年份,其表现仍与全数据训练模型高度竞争。结果表明,在固定计算预算下,广泛采样气候状态比时间连续性更具价值。该研究为区域气候降尺度与大规模集合投影工作流程提供了实用指导。
原文摘要 · Abstract (English)
High-resolution regional climate simulations provide critical information for climate impacts assessments but remain computationally expensive, motivating the development of machine-learning downscalers and emulators. A key challenge is determining how limited high-resolution simulations should be distributed across a changing climate trajectory to capture both forced climate response and internal variability. Using the CESM2 Large Ensemble over the western United States, we compare three training-year selection strategies under fixed data budgets: a contiguous block of historical years, years drawn from both the beginning and end of the simulation period, and years distributed throughout the full climate trajectory. Including both historical and future years consistently outperforms training on historical years alone, demonstrating the importance of exposing downscaling models to climate states outside the historical record and highlighting limitations of stationarity assumptions common in statistical downscaling. Training on years distributed throughout the full climate trajectory performs best overall, indicating that broad sampling of internal variability provides additional information beyond exposure to the forced climate response alone. Models trained on temporally distributed subsets more successfully reproduce variability in unseen ensemble members while retaining strong performance across a wide range of climate diagnostics. Even when trained on only one-tenth of the available high-resolution years, temporally distributed models remain highly competitive with full-data training. These results suggest that, under fixed computational budgets, broad sampling of climate states is more valuable than temporal continuity when allocating scarce high-resolution simulations. The findings provide practical guidance for regional climate downscaling and large-ensemble projection workflows.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。