arXiv:2510.15189cs.RO2025-10被引 2

无需人工示范,机器人可自主学习精准抓取操作。

RM-RL: Role-Model Reinforcement Learning for Precise Robot Manipulation

  • 用近优动作自动生成训练标签,替代人工示范。
  • 实测提升53%位移精度和20%旋转精度。
  • 适合需要高精度操作的实验室机器人场景。

精确机器人操作对化学与生物实验等精细任务至关重要,微小误差(如试剂溢出)可能导致任务失败。现有方法多依赖预收集的专家示范,通过模仿学习或离线强化学习训练策略,但高质量示范获取困难,且离线RL常面临分布偏移和数据效率低的问题。本文提出角色模型强化学习(RM-RL)框架,将在线与离线训练统一于真实环境。核心是角色模型策略:利用近似最优动作自动为在线数据生成标签,无需人工示范。RM-RL将策略学习转化为监督训练,缓解分布不匹配带来的不稳定性,提升训练效率。混合训练机制进一步将在线角色模型数据用于离线重用,通过重复采样增强数据利用效率。大量实验表明,RM-RL收敛更快更稳定,显著提升真实世界操作表现:位移精度提升53%,旋转精度提升20%。最终成功完成将细胞板精准放置到货架的挑战性任务,验证了该框架在以往方法失效场景下的有效性。

原文摘要 · Abstract (English)

Precise robot manipulation is critical for fine-grained applications such as chemical and biological experiments, where even small errors (e.g., reagent spillage) can invalidate an entire task. Existing approaches often rely on pre-collected expert demonstrations and train policies via imitation learning (IL) or offline reinforcement learning (RL). However, obtaining high-quality demonstrations for precision tasks is difficult and time-consuming, while offline RL commonly suffers from distribution shifts and low data efficiency. We introduce a Role-Model Reinforcement Learning (RM-RL) framework that unifies online and offline training in real-world environments. The key idea is a role-model strategy that automatically generates labels for online training data using approximately optimal actions, eliminating the need for human demonstrations. RM-RL reformulates policy learning as supervised training, reducing instability from distribution mismatch and improving efficiency. A hybrid training scheme further leverages online role-model data for offline reuse, enhancing data efficiency through repeated sampling. Extensive experiments show that RM-RL converges faster and more stably than existing RL methods, yielding significant gains in real-world manipulation: 53% improvement in translation accuracy and 20% in rotation accuracy. Finally, we demonstrate the successful execution of a challenging task, precisely placing a cell plate onto a shelf, highlighting the framework's effectiveness where prior methods fail.

机器人操控强化学习自监督精准操作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。