arXiv:2410.23737cs.LG2024-10

提出非一体式策略方法,协调离线策略的利用与在线策略的探索。

A Non-Monolithic Policy Approach of Offline-to-Online Reinforcement Learning

  • 分离离线策略的利用与在线策略的探索,不修改离线策略
  • 相比PEX方法,性能更优,提升数据效率与收敛速度
  • 适合需快速适应新任务的强化学习场景

离线到在线强化学习结合预训练的离线策略与针对下游任务训练的在线策略,旨在提升数据效率并加速性能提升。现有策略扩展(PEX)方法使用由两类策略组成的策略集,但未修改离线策略进行探索和学习,导致在线策略学习不足。由于预训练离线策略基于先验经验可辅助在线策略在下游任务中有效利用,应针对性地发挥其优势;而行为策略尚不成熟的在线策略则具备探索潜力。因此,本研究聚焦于协调离线策略的利用与在线策略的探索,提出一种非一体式探索的离线到在线强化学习方法。实验表明,该方法在性能上优于现有PEX方法。

原文摘要 · Abstract (English)

Offline-to-online reinforcement learning (RL) leverages both pre-trained offline policies and online policies trained for downstream tasks, aiming to improve data efficiency and accelerate performance enhancement. An existing approach, Policy Expansion (PEX), utilizes a policy set composed of both policies without modifying the offline policy for exploration and learning. However, this approach fails to ensure sufficient learning of the online policy due to an excessive focus on exploration with both policies. Since the pre-trained offline policy can assist the online policy in exploiting a downstream task based on its prior experience, it should be executed effectively and tailored to the specific requirements of the downstream task. In contrast, the online policy, with its immature behavioral strategy, has the potential for exploration during the training phase. Therefore, our research focuses on harmonizing the advantages of the offline policy, termed exploitation, with those of the online policy, referred to as exploration, without modifying the offline policy. In this study, we propose an innovative offline-to-online RL method that employs a non-monolithic exploration approach. Our methodology demonstrates superior performance compared to PEX.

强化学习离线-在线策略协同

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。