arXiv:2510.19528stat.MLcs.LG2025-10

用离线数据学习价值上下界,加速在线强化学习。

Learning Upper Lower Value Envelopes to Shape Online RL: A Principled Approach

  • 分两阶段:先从离线数据学价值上下界,再用于在线训练
  • 在表格MDP上显著降低后悔值,优于UCBVI和已有方法
  • 上下界为数据驱动且随机变量建模,理论更可信

我们研究利用离线数据加速在线强化学习的根本问题——这一方向潜力巨大但理论基础薄弱。本文聚焦于如何学习并应用价值包络。提出一种原理严谨的两阶段框架:第一阶段利用离线数据推导价值函数的上下界;第二阶段将这些边界融入在线算法。相比以往方法,本工作解耦上下界,实现更灵活、更紧致的近似。与依赖固定形貌函数的方法不同,我们的包络是数据驱动的,并显式建模为随机变量,通过滤化论证保证各阶段独立性。理论分析建立了由两个可解释量决定的高概率后悔界,正式连接了离线预训练与在线微调。在表格MDP上的实验表明,相比UCBVI及先前方法,后悔值大幅下降,同时保持与相关方法相当的竞争力。

原文摘要 · Abstract (English)

We investigate the fundamental problem of leveraging offline data to accelerate online reinforcement learning - a direction with strong potential but limited theoretical grounding. Our study centers on how to \emph{learn} and \emph{apply} value envelopes within this context. To this end, we introduce a principled two-stage framework: the first stage uses offline data to derive upper and lower bounds on value functions, while the second incorporates these learned bounds into online algorithms. Our method extends prior work by decoupling the upper and lower bounds, enabling more flexible and tighter approximations. In contrast to approaches that rely on fixed shaping functions, our envelopes are data-driven and explicitly modeled as random variables, with a filtration argument ensuring independence across phases. The analysis establishes high-probability regret bounds determined by two interpretable quantities, thereby providing a formal bridge between offline pre-training and online fine-tuning. Empirical results on tabular MDPs demonstrate substantial regret reductions compared with both UCBVI and prior methods while remaining competitive with related approaches.

强化学习离线学习价值函数

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。