用强化学习让模型‘预演’未来,提升时空预测准确性。
Spatiotemporal Forecasting as Planning: A Model-Based Reinforcement Learning Approach with Generative World Models
- 将预测任务转化为规划问题,用生成世界模型模拟多种未来场景。
- 通过非可微指标作为奖励,找到高回报的预测序列,误差显著降低。
- 特别适合需要捕捉极端事件的时空预测任务,如气象、交通建模。
为应对物理时空预测中固有的随机性与不可微指标的双重挑战,我们提出时空预测即规划(SFP),一种基于模型强化学习的新范式。SFP构建了一个新型生成世界模型,用于模拟多样且高保真的未来状态,实现‘基于想象’的环境仿真。在此框架下,基础预测模型作为智能体,由基于束搜索的规划算法引导,利用不可微领域指标作为奖励信号,探索高回报的未来序列。这些高奖励候选序列随后作为伪标签,通过迭代自训练持续优化智能体策略,显著降低预测误差,并在捕捉极端事件等关键领域指标上表现卓越。
原文摘要 · Abstract (English)
To address the dual challenges of inherent stochasticity and non-differentiable metrics in physical spatiotemporal forecasting, we propose Spatiotemporal Forecasting as Planning (SFP), a new paradigm grounded in Model-Based Reinforcement Learning. SFP constructs a novel Generative World Model to simulate diverse, high-fidelity future states, enabling an "imagination-based" environmental simulation. Within this framework, a base forecasting model acts as an agent, guided by a beam search-based planning algorithm that leverages non-differentiable domain metrics as reward signals to explore high-return future sequences. These identified high-reward candidates then serve as pseudo-labels to continuously optimize the agent's policy through iterative self-training, significantly reducing prediction error and demonstrating exceptional performance on critical domain metrics like capturing extreme events.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。