用自生成课程的强化学习,让机器人精准操控柔性长条物体。
Self-Curriculum Model-based Reinforcement Learning for Shape Control of Deformable Linear Objects
- 分两阶段:大变形用多模型强化学习,小变形用视觉伺服精调。
- 仿真中成功率显著超越主流方法,30个不同任务全成功。
- 无需额外训练即可直接用于真实世界,适应多种材料和形状。
精确控制柔性线性物体(DLOs)在工业与医疗等机器人应用中至关重要。现有方法难以处理含相反曲率等复杂大变形任务,且效率与精度不足。为此,我们提出一种两阶段框架,结合强化学习(RL)与在线视觉伺服。在大变形阶段,采用基于集成动力学模型的模型化强化学习,大幅提升样本效率;并设计自生成课程的目标机制,通过想象评估动态选择高多样性、中等难度目标,优化策略学习。在小变形阶段,部署基于雅可比矩阵的视觉伺服控制器,确保高精度收敛。仿真结果表明,该方法实现高效策略学习,在形状控制成功率与精度上显著优于主流基线。此外,训练好的策略可零样本迁移至真实任务,成功完成30个涵盖不同尺寸与材质的DLO任务,覆盖多样化初始与目标形态。
原文摘要 · Abstract (English)
Precise shape control of Deformable Linear Objects (DLOs) is crucial in robotic applications such as industrial and medical fields. However, existing methods face challenges in handling complex large deformation tasks, especially those involving opposite curvatures, and lack efficiency and precision. To address this, we propose a two-stage framework combining Reinforcement Learning (RL) and online visual servoing. In the large-deformation stage, a model-based reinforcement learning approach using an ensemble of dynamics models is introduced to significantly improve sample efficiency. Additionally, we design a self-curriculum goal generation mechanism that dynamically selects intermediate-difficulty goals with high diversity through imagined evaluations, thereby optimizing the policy learning process. In the small-deformation stage, a Jacobian-based visual servo controller is deployed to ensure high-precision convergence. Simulation results show that the proposed method enables efficient policy learning and significantly outperforms mainstream baselines in shape control success rate and precision. Furthermore, the framework effectively transfers the policy trained in simulation to real-world tasks with zero-shot adaptation. It successfully completes all 30 cases with diverse initial and target shapes across DLOs of different sizes and materials. The project website is available at: https://anonymous.4open.science/w/sc-mbrl-dlo-EB48/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。