arXiv:2505.08078cs.ROcs.AI2025-05被引 15

提出机器人批量在线强化学习有效方法,提升自收集数据利用效率。

What Matters for Batch Online Reinforcement Learning in Robotics?

  • 用Q函数引导学习,优于传统模仿学习方法
  • 通过策略分布选最优动作,提升性能与收敛性
  • 高表达力策略配合时序噪声,显著改善学习效果

从自主收集的大量数据中进行策略优化的能力——我们称之为批量在线强化学习——有望通过大幅减少人工数据收集需求,实现真正可扩展的机器人学习。然而,由于算法难以有效利用自主数据,该范式仍具挑战性。现有方法如模仿学习和过滤模仿学习常无法高效改进或快速收敛至次优解。为此,我们系统研究了三个维度:(i)算法类别、(ii)策略提取方法、(iii)策略表达能力,分析其对性能及数据量扩展的影响。结果表明,使用Q函数指导学习显著优于模仿类方法;采用从策略分布中选择最优动作的隐式提取方式,优于传统离线强化学习策略提取;更具表达力的策略类表现更优。基于此,我们提出一套通用有效方案,并引入时序相关噪声以增强多样性,进一步提升性能。相比已有方法,本方案在性能与数据扩展性上均有显著提升。

原文摘要 · Abstract (English)

The ability to learn from large batches of autonomously collected data for policy improvement -- a paradigm we refer to as batch online reinforcement learning -- holds the promise of enabling truly scalable robot learning by significantly reducing the need for human effort of data collection while getting benefits from self-improvement. Yet, despite the promise of this paradigm, it remains challenging to achieve due to algorithms not being able to learn effectively from the autonomous data. For example, prior works have applied imitation learning and filtered imitation learning methods to the batch online RL problem, but these algorithms often fail to efficiently improve from the autonomously collected data or converge quickly to a suboptimal point. This raises the question of what matters for effective batch online RL in robotics. Motivated by this question, we perform a systematic empirical study of three axes -- (i) algorithm class, (ii) policy extraction methods, and (iii) policy expressivity -- and analyze how these axes affect performance and scaling with the amount of autonomous data. Through our analysis, we make several observations. First, we observe that the use of Q-functions to guide batch online RL significantly improves performance over imitation-based methods. Building on this, we show that an implicit method of policy extraction -- via choosing the best action in the distribution of the policy -- is necessary over traditional policy extraction methods from offline RL. Next, we show that an expressive policy class is preferred over less expressive policy classes. Based on this analysis, we propose a general recipe for effective batch online RL. We then show a simple addition to the recipe of using temporally-correlated noise to obtain more diversity results in further performance gains. Our recipe obtains significantly better performance and scaling compared to prior methods.

强化学习机器人批量学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。