让机器人通过模仿成功轨迹提升控制效率
TDMPBC: Self-Imitative Reinforcement Learning for Humanoid Robot Control
- 引入自模仿机制,动态强化高回报轨迹的学习
- 在HumanoidBench上性能提升120%,仅多出5%计算开销
- 适合需要高效探索的复杂机器人控制任务
高自由度、复杂动作空间的人形机器人控制对强化学习算法构成挑战,需在有限样本下平衡探索与利用。在人形机器人运动控制中,大部分状态会导致摔倒,仅有极小部分能保持直立以完成任务。一旦探索到潜在有效区域,应更重视该区域数据。为此,我们提出自模仿强化学习(SIRL)框架,利用轨迹回报判断其任务相关性,并动态调整行为克隆权重。实验表明,该方法在挑战性HumanoidBench上实现120%性能提升,额外计算开销仅为5%。可视化分析显示,性能提升对应实际行为改进,多个任务得以成功解决。
原文摘要 · Abstract (English)
Complex high-dimensional spaces with high Degree-of-Freedom and complicated action spaces, such as humanoid robots equipped with dexterous hands, pose significant challenges for reinforcement learning (RL) algorithms, which need to wisely balance exploration and exploitation under limited sample budgets. In general, feasible regions for accomplishing tasks within complex high-dimensional spaces are exceedingly narrow. For instance, in the context of humanoid robot motion control, the vast majority of space corresponds to falling, while only a minuscule fraction corresponds to standing upright, which is conducive to the completion of downstream tasks. Once the robot explores into a potentially task-relevant region, it should place greater emphasis on the data within that region. Building on this insight, we propose the $\textbf{S}$elf-$\textbf{I}$mitative $\textbf{R}$einforcement $\textbf{L}$earning ($\textbf{SIRL}$) framework, where the RL algorithm also imitates potentially task-relevant trajectories. Specifically, trajectory return is utilized to determine its relevance to the task and an additional behavior cloning is adopted whose weight is dynamically adjusted based on the trajectory return. As a result, our proposed algorithm achieves 120% performance improvement on the challenging HumanoidBench with 5% extra computation overhead. With further visualization, we find the significant performance gain does lead to meaningful behavior improvement that several tasks are solved successfully.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。