用能量感知框架提升机器人任务空间操作的安全性与稳定性
Contact-Safe Reinforcement Learning with ProMP Reparameterization and Energy Awareness
- 结合PPO与运动基元生成安全任务空间轨迹
- 在多种表面环境下实现高成功率与平滑运动
- 适合需要接触安全的复杂机械臂操作场景
基于马尔可夫决策过程的强化学习方法通常在机器人关节空间中应用,依赖有限的任务信息和对3D环境的部分感知。相比之下,基于回合的强化学习在轨迹一致性、任务感知和复杂任务表现方面具有优势。然而,传统逐步和回合式RL方法常忽略任务空间操作中的丰富接触信息,特别是在接触安全性和鲁棒性方面。本文提出一种面向接触丰富的操作任务的、基于任务空间的能量安全框架,通过近端策略优化(PPO)与运动基元相结合,生成可靠且安全的任务空间轨迹。此外,在该框架中引入能量感知的笛卡尔阻抗控制器目标,以确保机器人与环境之间的安全交互。实验结果表明,该框架在不同类型的3D表面环境中均优于现有方法,实现了高成功率、平滑轨迹以及能量安全的交互。
原文摘要 · Abstract (English)
Reinforcement learning (RL) approaches based on Markov Decision Processes (MDPs) are predominantly applied in the robot joint space, often relying on limited task-specific information and partial awareness of the 3D environment. In contrast, episodic RL has demonstrated advantages over traditional MDP-based methods in terms of trajectory consistency, task awareness, and overall performance in complex robotic tasks. Moreover, traditional step-wise and episodic RL methods often neglect the contact-rich information inherent in task-space manipulation, especially considering the contact-safety and robustness. In this work, contact-rich manipulation tasks are tackled using a task-space, energy-safe framework, where reliable and safe task-space trajectories are generated through the combination of Proximal Policy Optimization (PPO) and movement primitives. Furthermore, an energy-aware Cartesian Impedance Controller objective is incorporated within the proposed framework to ensure safe interactions between the robot and the environment. Our experimental results demonstrate that the proposed framework outperforms existing methods in handling tasks on various types of surfaces in 3D environments, achieving high success rates as well as smooth trajectories and energy-safe interactions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。