用飓风路径做输入反而让海洋模拟器更不准,因信号太罕见。
Too Rare to Learn: Prescribed Cyclone Tracks Degrade a Bay of Bengal Ocean Emulator

- 用预设飓风路径作为输入,但训练时只出现7.9%的天数
- 含路径输入的模型全输过期模型,误差集中在风暴区域
- 换掉真实路径图后,预测准确率提升7.5%~16.4%,适合气候建模者
神经海洋模拟器正被用于热带气旋频发海域的区域预报,自然设计是将飓风路径作为预设输入。我们在孟加拉湾测试此方法,发现其有害。从GLORYS12再分析数据中剔除15个完整飓风(强度65至150节),对比两个仅在四通道上不同(是否包含预设飓风路径)的U-Net。在三个随机种子下,仅含海洋输入的模型始终优于持续性模型,而条件模型则在所有运行中均落后,两模型技能范围无重叠(p = 3.1e-5,跨风暴配对检验)。原因在于暴露频率而非信号内容:路径通道仅在训练日7.9%的时间非零,一旦激活即超出分布。额外误差集中在预设风暴范围内;推理时用无风暴图替代真实路径图,对保留风暴的预测准确率提升7.5%至16.4%(每个种子均如此)。条件网络学到了对罕见信号的自信错误响应。
原文摘要 · Abstract (English)
Neural ocean emulators are being proposed for regional forecasting in cyclone-exposed coastal seas, and a natural design choice is to hand the network the cyclone as a prescribed input. We test that choice in the Bay of Bengal and find it harmful. We withhold 15 whole cyclones spanning 65 to 150 kt from GLORYS12 reanalysis and compare two U-Nets that are identical except for four prescribed cyclone-track channels. Across three seeds the ocean-only model beats persistence in every run and the storm-conditioned model loses to it in every run, with the two skill ranges disjoint (p = 3.1e-5, paired across storms). The cause is exposure frequency rather than signal content: the channels are non-zero on only 7.9% of training days, so they are out of distribution the moment they activate. The extra error falls inside the prescribed storm footprint, and replacing the real cyclone map with a no-storm map at inference improves held-out storm forecasts by 7.5 to 16.4% in every seed. The conditioned network has learned a response to a rare signal that is confidently wrong.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。