arXiv:2602.20220cs.ROcs.AI2026-02被引 2

实验证明,正确选择训练设计可让机器人在线强化学习更稳定。

What Matters for Simulation to Online Reinforcement Learning on Real Robots

  • 系统性测试100次真实机器人实验,对比不同设计影响
  • 发现常见默认设置可能有害,特定做法能稳定跨任务学习
  • 适合想快速部署在线强化学习的工程实践者

我们研究了哪些具体设计选择能促成物理机器人上在线强化学习的成功。在三种不同机器人平台上进行了100次真实世界训练,系统性地消融了算法、系统和实验决策等通常被隐含处理的因素。结果表明,一些广泛使用的默认设置可能有害,而一组稳健且易于采纳的标准强化学习实践则能在多种任务和硬件上实现稳定学习。这些发现是首个大规模实证研究此类设计选择的工作,使从业者能以更低工程成本部署在线强化学习。

原文摘要 · Abstract (English)

We investigate what specific design choices enable successful online reinforcement learning (RL) on physical robots. Across 100 real-world training runs on three distinct robotic platforms, we systematically ablate algorithmic, systems, and experimental decisions that are typically left implicit in prior work. We find that some widely used defaults can be harmful, while a set of robust, readily adopted design choices within standard RL practice yield stable learning across tasks and hardware. These results provide the first large-sample empirical study of such design choices, enabling practitioners to deploy online RL with lower engineering effort.

强化学习机器人在线学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。