预测时长越长,模型越难训练,但能更好泛化到短期预测。
Temporal horizons in forecasting: a performance-learnability trade-off
- 分析损失曲面几何随预测时长变化的规律。
- 混沌系统中损失粗糙度指数增长,周期系统线性增长。
- 长时训练模型对短期预测泛化更好,适合实际应用。
在训练自回归模型预测动态系统时,关键问题在于应预测多远的未来。过短的时长会遗漏长期趋势,过长则因累积预测误差阻碍收敛。本文通过分析损失曲面几何与训练时长的关系,揭示这一权衡。对于混沌系统,损失曲面粗糙度随训练时长呈指数增长;对于极限环系统,则呈线性增长,导致长时训练本质困难。然而,长时训练模型在短期预测上泛化性能优异,而短时训练模型在混沌系统中长期预测误差呈指数恶化,在周期系统中呈线性恶化。数值实验验证了理论,并为自回归预测模型的超参数选择提供理论依据。
原文摘要 · Abstract (English)
When training autoregressive models to forecast dynamical systems, a critical question arises: how far into the future should the model be trained to predict? Too short a horizon may miss long-term trends, while too long a horizon can impede convergence due to accumulating prediction errors. In this work, we formalize this trade-off by analyzing how the geometry of the loss landscape depends on the training horizon. We prove that for chaotic systems, the loss landscape's roughness grows exponentially with the training horizon, while for limit cycles, it grows linearly, making long-horizon training inherently challenging. However, we also show that models trained on long horizons generalize well to short-term forecasts, whereas those trained on short horizons suffer exponentially (resp. linearly) worse long-term predictions in chaotic (resp. periodic) systems. We validate our theory through numerical experiments and discuss practical implications for selecting training horizons. Our results provide a principled foundation for hyperparameter optimization in autoregressive forecasting models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。