arXiv:2409.20568cs.ROcs.AI2024-09CoRL被引 32

无需人工干预,机器人在真实世界中持续提升移动操作能力。

Continuously Improving Mobile Manipulation with Autonomous Real-World RL

  • 通过任务相关自主性引导探索,避免在目标附近停滞。
  • 利用行为先验实现高效策略学习,平均成功率80%。
  • 结合语义信息与细粒度观测构建通用奖励函数,适合工业部署。

我们提出一种完全自主的真实世界强化学习框架,用于移动操作任务,可在无大量传感器或人工监督的情况下学习策略。该框架基于三项设计:1)任务相关的自主性,引导探索聚焦于物体交互,防止在目标状态附近停滞;2)利用基础任务知识构建行为先验,实现高效策略学习;3)构建通用奖励函数,融合人类可理解的语义信息与低层细粒度观测。我们在Spot机器人上验证了该方法在四项挑战性移动操作任务上的持续改进能力,平均成功率达到80%,较现有方法提升3-4倍。视频展示见 https://continual-mobile-manip.github.io/

原文摘要 · Abstract (English)

We present a fully autonomous real-world RL framework for mobile manipulation that can learn policies without extensive instrumentation or human supervision. This is enabled by 1) task-relevant autonomy, which guides exploration towards object interactions and prevents stagnation near goal states, 2) efficient policy learning by leveraging basic task knowledge in behavior priors, and 3) formulating generic rewards that combine human-interpretable semantic information with low-level, fine-grained observations. We demonstrate that our approach allows Spot robots to continually improve their performance on a set of four challenging mobile manipulation tasks, obtaining an average success rate of 80% across tasks, a 3-4 improvement over existing approaches. Videos can be found at https://continual-mobile-manip.github.io/

强化学习机器人操作自主学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。