arXiv:2512.01987cs.LGcs.AI2025-12NeurIPS

提出新框架FORL,让离线强化学习适应真实世界动态变化环境。

Forecasting in Offline Reinforcement Learning for Non-stationary Environments

  • 用条件扩散模型生成未来状态候选,不依赖特定非平稳模式。
  • 在引入真实时间序列数据的基准上,性能显著优于现有方法。
  • 适合需即时应对突发环境变化的工业级强化学习应用。

离线强化学习(Offline RL)在无法获取额外交互数据时,提供从预收集数据集训练策略的可行路径。然而,现有方法常假设环境平稳或仅在测试时考虑合成扰动,这在实际中往往失效,因真实场景存在突发、随时间变化的偏移,导致部分可观测性,使智能体误判自身状态并降低性能。为此,我们提出非平稳离线强化学习中的预测框架(FORL),统一了(i)无需预设未来非平稳模式的条件扩散候选状态生成,以及(ii)零样本时间序列基础模型。FORL针对可能非马尔可夫的意外偏移环境,要求智能体在每轮开始即具备鲁棒性能。在加入真实时间序列数据以模拟现实非平稳性的离线RL基准上,实证表明FORL持续优于竞争基线。通过将零样本预测与智能体经验融合,旨在弥合离线RL与真实复杂非平稳环境之间的差距。

原文摘要 · Abstract (English)

Offline Reinforcement Learning (RL) provides a promising avenue for training policies from pre-collected datasets when gathering additional interaction data is infeasible. However, existing offline RL methods often assume stationarity or only consider synthetic perturbations at test time, assumptions that often fail in real-world scenarios characterized by abrupt, time-varying offsets. These offsets can lead to partial observability, causing agents to misperceive their true state and degrade performance. To overcome this challenge, we introduce Forecasting in Non-stationary Offline RL (FORL), a framework that unifies (i) conditional diffusion-based candidate state generation, trained without presupposing any specific pattern of future non-stationarity, and (ii) zero-shot time-series foundation models. FORL targets environments prone to unexpected, potentially non-Markovian offsets, requiring robust agent performance from the onset of each episode. Empirical evaluations on offline RL benchmarks, augmented with real-world time-series data to simulate realistic non-stationarity, demonstrate that FORL consistently improves performance compared to competitive baselines. By integrating zero-shot forecasting with the agent's experience, we aim to bridge the gap between offline RL and the complexities of real-world, non-stationary environments.

离线RL非平稳环境时间序列扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。