对比6种时序大模型在加州野火PM2.5预测中的泛化能力,发现小模型仍更可靠。
Evaluating the Generalizability of Foundation Models for Extreme Environmental Events: Case Study of California Wildfire PM2.5

- 用留一事件外验证法测试大模型对未见野火的预测能力
- 小模型BiLSTM在极端值预测上准确率最高,达0.63
- 微调可修复大模型稳定性问题,但仍未超越传统模型
野火烟雾导致极端高浓度PM₂.₅,严重威胁公共健康,但预测罕见危险级峰值仍是根本挑战。时间序列基础模型(TSFMs)虽在通用基准表现优异,但在极端分布外条件下的行为尚不明确。本文首次系统性对比六种TSFM配置(零样本TimesFM、Chronos-2、Moirai-2、Time-MoE,以及LoRA微调后的Chronos-2和Time-MoE)与全训练基线(LSTM、BiLSTM、Transformer)及简单持续性模型,在覆盖79个监测站点、1,375次野火事件的12年(2013–2025)小时级PM₂.₅数据集上的表现。采用留一事件外(LOIO)协议,评估在6、12、24小时预测前视下对未见火灾的泛化能力,使用MAE、RMSE及美国环保署空气质量指数(AQI)阈值处的超出事件F1评分。结果表明,性能存在稳定层级:BiLSTM在所有指标中最低的MAE(5.16 μg/m³)和最高的超出事件F1(危险带>225.5 μg/m³时达0.63),优于任一基础模型;零样本TSFM仅小幅优于持续性模型,其中零样本Chronos-2在尾部呈现严重RMSE不稳定(23.4 μg/m³,负R²);LoRA微调显著提升适应性并修复此不稳定性,但无一基础模型在任何指标上超越训练好的循环基线。研究挑战了大模型在环境预测中普遍占优的假设,为野火空气质量预测提供可操作部署指导。
原文摘要 · Abstract (English)
Wildfire smoke events produce extreme PM$_{2.5}$ concentrations that pose severe public health risks, yet forecasting rare, hazardous-level spikes remains a fundamental challenge. Time series foundation models (TSFMs), pretrained models offering zero-shot inference and efficient adaptation, perform strongly on general benchmarks, but their behavior under extreme out-of-distribution conditions is poorly understood. We present the first systematic benchmark comparing six TSFM configurations (zero-shot TimesFM, Chronos-2, Moirai-2, and Time-MoE, plus LoRA fine-tuned Chronos-2 and Time-MoE) against fully-trained baselines (LSTM, BiLSTM, Transformer) and naive persistence on a 12-year (2013--2025) hourly PM$_{2.5}$ dataset covering 1,375 wildfire incidents across 79 California monitoring sites. A leave-one-incident-out (LOIO) protocol evaluates generalization to unseen fires, using MAE, RMSE, and exceedance F1 at EPA AQI thresholds across 6-, 12-, and 24-hour horizons. Results reveal a consistent hierarchy. The BiLSTM achieves the lowest MAE ($5.16\,μg/m^3$) and the highest exceedance F1 at every threshold, including the Hazardous band ($>225.5\,μg/m^3$), reaching 0.63 versus at most 0.54 for any foundation model. Zero-shot TSFMs improve on persistence only modestly, and zero-shot Chronos-2 exhibits severe RMSE tail instability ($23.4\,μg/m^3$, negative $R^2$) from sporadic large errors. LoRA fine-tuning substantially improves both adapted families and largely repairs this instability, yet no foundation model surpasses the trained recurrent baselines on any metric. These findings challenge the assumption that larger pretrained models universally dominate environmental forecasting and provide actionable deployment guidance for wildfire air quality prediction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。