研究气象预测的缩放规律,发现不同通道和时段表现差异大。
Towards Scaling Law Analysis For Spatiotemporal Weather Data
- 从单步损失扩展到长时滚动预测与分通道评估
- 全局误差缩放看似良好,但多数通道在远期预报中恶化
- 适合关注气象模型优化与资源分配的研究者
计算最优缩放规律在NLP和CV中研究较充分,其目标通常为单步且输出同质。而气象预测更复杂:自回归滚动会累积误差,多物理通道耦合且尺度与可预测性各异,全局测试指标与短时训练所暗示的局部、晚期预报行为常不一致。本文将神经网络缩放分析从单步训练损失扩展至长时滚动与分通道指标。我们量化了(1)预测误差在各通道的分布及随预报时长的增长速率;(2)当误差全局聚合时,测试误差是否符合幂律缩放关系;(3)该拟合在参数、数据、算力等不同缩放轴上如何随时长与通道变化。结果表明存在显著跨通道与跨时长异质性:全局缩放可能表现良好,但多数通道在远期预报中明显退化。本文讨论了加权目标、时长感知训练策略以及输出间资源分配的启示。
原文摘要 · Abstract (English)
Compute-optimal scaling laws are relatively well studied for NLP and CV, where objectives are typically single-step and targets are comparatively homogeneous. Weather forecasting is harder to characterize in the same framework: autoregressive rollouts compound errors over long horizons, outputs couple many physical channels with disparate scales and predictability, and globally pooled test metrics can disagree sharply with per-channel, late-lead behavior implied by short-horizon training. We extend neural scaling analysis for autoregressive weather forecasting from single-step training loss to long rollouts and per-channel metrics. We quantify (1) how prediction error is distributed across channels and how its growth rate evolves with forecast horizon, (2) if power law scaling holds for test error, relative to rollout length when error is pooled globally, and (3) how that fit varies jointly with horizon and channel for parameter, data, and compute-based scaling axes. We find strong cross-channel and cross-horizon heterogeneity: pooled scaling can look favorable while many channels degrade at late leads. We discuss implications for weighted objectives, horizon-aware curricula, and resource allocation across outputs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。