arXiv:2512.00383cs.LGcs.AI2025-12

把离线强化学习当作在线学习的子程序,提升学习效率。

An Empirical Study on the Effectiveness of Incorporating Offline RL As Online RL Subroutines

  • 用历史交互数据作为离线数据,驱动在线学习
  • 新方法显著提升在线学习效率,而旧方法效果差
  • 适合需要高效探索的复杂任务研究者

本文提出将离线强化学习算法作为从零开始的在线强化学习的子程序。这一思路可行,因为在线学习代理可将其历史交互数据重新利用为离线数据集。我们构建了一个框架,支持多种离线学习集成方式,如最终策略推荐和在线微调。同时引入若干便捷技术以提升其有效性。大量系统性实证分析表明:1)该框架的有效性高度依赖任务特性;2)所提技术显著增强其性能;3)现有在线微调方法整体无效,亟需进一步研究。

原文摘要 · Abstract (English)

We take the novel perspective of incorporating offline RL algorithms as subroutines of tabula rasa online RL. This is feasible because an online learning agent can repurpose its historical interactions as offline dataset. We formalize this idea into a framework that accommodates several variants of offline RL incorporation such as final policy recommendation and online fine-tuning. We further introduce convenient techniques to improve its effectiveness in enhancing online learning efficiency. Our extensive and systematic empirical analyses show that 1) the effectiveness of the proposed framework depends strongly on the nature of the task, 2) our proposed techniques greatly enhance its effectiveness, and 3) existing online fine-tuning methods are overall ineffective, calling for more research therein.

强化学习离线学习在线学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。