arXiv:2603.20315cs.LG2026-03被引 1

滚动验证发现:传统方法反而比机器学习更稳定可靠

Rolling-Origin Validation Reverses Model Rankings in Multi-Step PM10 Forecasting: XGBoost, SARIMA, and Persistence

  • 采用滚动起源验证,每月更新模型评估
  • XGBoost在短期预测中不如简单基准,SARIMA持续有效
  • 适合关注实际应用性能的研究者与环保从业者

许多空气质量预测研究宣称机器学习带来提升,但评估常使用静态时间划分且忽略基准方法,难以反映日常更新下的实际价值。基于2017至2024年南欧城市背景站的2,350条日均PM10观测数据,我们比较了XGBoost与SARIMA在静态划分和每月更新的滚动起源协议下对1至7天前的预测表现。结果显示,静态评估显示XGBoost在1-7天内表现良好,但滚动起源评估反转排名:在短中期预测中XGBoost并不持续优于持久性基准;而SARIMA在整个预测范围保持正向相对技能。对于研究者,静态划分可能夸大模型的实际效用并改变排名;对于实践者,滚动起源与持久性参照的技能曲线能清晰揭示各方法在不同预报时长下的可靠性。

原文摘要 · Abstract (English)

(a) Many air quality forecasting studies report gains from machine learning, but evaluations often use static chronological splits and omit persistence baselines, so the operational added value under routine updating is unclear. (b) Using 2,350 daily PM10 observations from 2017 to 2024 at an urban background monitoring station in southern Europe, we compare XGBoost and SARIMA against persistence under a static split and a rolling-origin protocol with monthly updates. We report horizon-specific skill and the predictability horizon, defined as the maximum horizon with positive persistence-relative skill. Static evaluation suggests XGBoost performs well from one to seven days ahead, but rolling-origin evaluation reverses rankings: XGBoost is not consistently better than persistence at short and intermediate horizons, whereas SARIMA remains positively skilled across the full range. (c) For researchers, static splits can overstate operational usefulness and change rankings. For practitioners, rolling-origin, persistence-referenced skill profiles show which methods stay reliable at each lead time.

空气污染预测滚动验证时间序列模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。