用强化学习快速生成机械臂路径跟踪的优质初始轨迹
Learning-based Initialization of Trajectory Optimization for Path-following Problems of Redundant Manipulators
- 基于专家示范的强化学习生成初始轨迹
- 在仿真与真实机械臂上均提升优化效率与精度
- 适合需要快速生成可靠轨迹的工业场景
轨迹优化(TO)是生成冗余机械臂沿六维笛卡尔路径运动关节轨迹的有效工具。优化性能很大程度上依赖于初始轨迹的质量。然而,由于解空间极其庞大且配置空间中缺乏任务约束先验知识,高质量初始轨迹的选择非常困难,耗时较长。为缓解此问题,我们提出一种基于学习的初始轨迹生成方法,通过示例引导的强化学习,在短时间内生成高质量初始轨迹。此外,我们设计了一种零空间投影的模仿奖励,通过高效学习专家示范中捕捉到的运动学可行性,考虑零空间约束。统计评估显示,将本方法输出作为输入后,相比三种基线方法,轨迹优化在最优性、效率和适用性方面均有提升。真实世界实验也验证了七自由度机械臂上的性能改进与可行性。
原文摘要 · Abstract (English)
Trajectory optimization (TO) is an efficient tool to generate a redundant manipulator's joint trajectory following a 6-dimensional Cartesian path. The optimization performance largely depends on the quality of initial trajectories. However, the selection of a high-quality initial trajectory is non-trivial and requires a considerable time budget due to the extremely large space of the solution trajectories and the lack of prior knowledge about task constraints in configuration space. To alleviate the issue, we present a learning-based initial trajectory generation method that generates high-quality initial trajectories in a short time budget by adopting example-guided reinforcement learning. In addition, we suggest a null-space projected imitation reward to consider null-space constraints by efficiently learning kinematically feasible motion captured in expert demonstrations. Our statistical evaluation in simulation shows the improved optimality, efficiency, and applicability of TO when we plug in our method's output, compared with three other baselines. We also show the performance improvement and feasibility via real-world experiments with a seven-degree-of-freedom manipulator.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。