arXiv:2502.05595cs.RO2025-02被引 3

用模型引导强化学习,让机器人高效精准投掷物体

Data efficient Robotic Object Throwing with Model-Based Reinforcement Learning

  • 基于模型的强化学习框架,融合数据驱动建模与策略优化
  • 仅需少量交互数据即可完成训练,在仿真和真实机器人上均表现优异
  • 适合需要快速适应新目标的工业投掷场景

抓取与放置(PnP)是工业机器人中的基础操作,但受限于工作空间,灵活性不足。抓取并投掷(PnT)通过利用重力等外部资源,可提升效率并扩大作业范围。然而,其执行复杂,需精确协调高速运动与物体动力学。现有方法分为解析式与基于学习的方法:解析法建模耗时且泛化能力弱;基于模型无关强化学习(MFRL)虽具自动化与适应性,但需大量交互。本文提出基于模型的强化学习框架MC-PILOT,结合数据驱动建模与策略优化,有效处理模型不确定性与释放误差。在仿真与真实世界测试中,使用Franka Emika Panda机械臂验证,该方法能快速泛化至新目标,性能优于解析法与模型无关方法。

原文摘要 · Abstract (English)

Pick-and-place (PnP) operations, featuring object grasping and trajectory planning, are fundamental in industrial robotics applications. Despite many advancements in the field, PnP is limited by workspace constraints, reducing flexibility. Pick-and-throw (PnT) is a promising alternative where the robot throws objects to target locations, leveraging extrinsic resources like gravity to improve efficiency and expand the workspace. However, PnT execution is complex, requiring precise coordination of high-speed movements and object dynamics. Solutions to the PnT problem are categorized into analytical and learning-based approaches. Analytical methods focus on system modeling and trajectory generation but are time-consuming and offer limited generalization. Learning-based solutions, in particular Model-Free Reinforcement Learning (MFRL), offer automation and adaptability but require extensive interaction time. This paper introduces a Model-Based Reinforcement Learning (MBRL) framework, MC-PILOT, which combines data-driven modeling with policy optimization for efficient and accurate PnT tasks. MC-PILOT accounts for model uncertainties and release errors, demonstrating superior performance in simulations and real-world tests with a Franka Emika Panda manipulator. The proposed approach generalizes rapidly to new targets, offering advantages over analytical and Model-Free methods.

机器人控制强化学习模型预测工业自动化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。