用深度强化学习实现通用低成本避障机械臂规划
URPlanner: A Universal Paradigm For Collision-Free Robotic Motion Planning Based on Deep Reinforcement Learning
- 设计参数化任务空间与无距离依赖的避障奖励
- 提出增强探索算法,提升多种DRL模型性能
- 仅需少量专家示范即可生成大规模轨迹数据
在复杂环境中对冗余机械臂进行无碰撞运动规划仍待深入探索。尽管深度强化学习(DRL)与机器人结合展现出处理多样化任务的潜力,但现有基于DRL的机械臂避障规划方法成本高昂,限制了实际部署。原因在于过度依赖最小距离、DRL探索与决策能力不足,以及数据获取与利用效率低下。本文提出URPlanner,一种基于DRL的通用无碰撞机械臂运动规划范式。该方法具备平台无关性、训练与部署成本低、适用于任意机械臂且无需求解逆运动学等优势。首先,构建参数化任务空间与独立于最小距离的通用避障奖励函数;其次,引入增强策略探索与评估算法,可适配多种DRL算法以提升性能;第三,提出专家数据扩散策略,仅需少量专家示范即可生成大规模轨迹数据集。实验全面验证了所提方法的优越性。
原文摘要 · Abstract (English)
Collision-free motion planning for redundant robot manipulators in complex environments is yet to be explored. Although recent advancements at the intersection of deep reinforcement learning (DRL) and robotics have highlighted its potential to handle versatile robotic tasks, current DRL-based collision-free motion planners for manipulators are highly costly, hindering their deployment and application. This is due to an overreliance on the minimum distance between the manipulator and obstacles, inadequate exploration and decision-making by DRL, and inefficient data acquisition and utilization. In this article, we propose URPlanner, a universal paradigm for collision-free robotic motion planning based on DRL. URPlanner offers several advantages over existing approaches: it is platform-agnostic, cost-effective in both training and deployment, and applicable to arbitrary manipulators without solving inverse kinematics. To achieve this, we first develop a parameterized task space and a universal obstacle avoidance reward that is independent of minimum distance. Second, we introduce an augmented policy exploration and evaluation algorithm that can be applied to various DRL algorithms to enhance their performance. Third, we propose an expert data diffusion strategy for efficient policy learning, which can produce a large-scale trajectory dataset from only a few expert demonstrations. Finally, the superiority of the proposed methods is comprehensively verified through experiments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。