arXiv:2411.03810cs.LGstat.ML2024-11被引 12

利用带动态偏移的历史数据,提升强化学习采样效率。

Hybrid Transfer Reinforcement Learning: Provable Sample Efficiency from Shifted-Dynamics Data

  • 提出混合迁移强化学习框架,结合目标环境在线数据与源环境离线数据。
  • 在已知动态偏移程度时,算法实现问题相关采样复杂度,优于纯在线学习。
  • 实验验证该方法显著超越现有在线强化学习基线,适合有历史数据的场景。

在线强化学习通常需要大量高风险的在线交互数据来学习目标任务策略,这促使研究者关注如何利用历史数据提升采样效率。历史数据可能来自动态不同的旧环境或相关环境。当前尚不清楚如何有效利用此类数据以可证明的方式提升学习效率。为此,我们提出了混合迁移强化学习(HTRL)设置:智能体在目标环境中学习,同时访问源环境的离线数据,而源环境具有动态偏移。我们证明,在缺乏动态偏移信息的情况下,一般性的偏移动态数据,即使偏移微小,也无法降低目标环境中的采样复杂度。然而,若事先知晓动态偏移的程度,我们设计了HySRL算法,实现了问题相关的采样复杂度,并优于纯在线强化学习。实验结果表明,HySRL显著超越现有的在线强化学习基线。

原文摘要 · Abstract (English)

Online Reinforcement learning (RL) typically requires high-stakes online interaction data to learn a policy for a target task. This prompts interest in leveraging historical data to improve sample efficiency. The historical data may come from outdated or related source environments with different dynamics. It remains unclear how to effectively use such data in the target task to provably enhance learning and sample efficiency. To address this, we propose a hybrid transfer RL (HTRL) setting, where an agent learns in a target environment while accessing offline data from a source environment with shifted dynamics. We show that -- without information on the dynamics shift -- general shifted-dynamics data, even with subtle shifts, does not reduce sample complexity in the target environment. However, with prior information on the degree of the dynamics shift, we design HySRL, a transfer algorithm that achieves problem-dependent sample complexity and outperforms pure online RL. Finally, our experimental results demonstrate that HySRL surpasses state-of-the-art online RL baseline.

强化学习迁移学习样本效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。