用基础模型提升人形机器人行走与操作的稳定性。
FLAM: Foundation Model-Based Body Stabilization for Humanoid Locomotion and Manipulation
- 引入姿态稳定奖励,结合虚拟人体重建优化动作
- 在基准测试中性能超越现有先进RL方法
- 适合研究人形机器人控制与强化学习的学者
近年来,人形机器人受到广泛关注。强化学习(RL)是控制人形机器人全身的主要方法之一,通过环境交互和任务奖励来学习行为。然而,现有RL方法很少显式考虑身体稳定性对行走与操作的影响。仅依赖任务奖励的RL方法在实现全身控制高性能方面仍面临挑战。本文提出一种基于基础模型的人形机器人行走与操作方法(FLAM)。FLAM将一个稳定奖励函数与基础策略结合:首先将机器人姿态映射到3D虚拟人体模型;然后通过人体运动重建模型进行姿态稳定与重构;最后利用重构前后的姿态差计算稳定奖励。将该奖励与任务奖励结合后,有效引导策略学习。在人形机器人基准测试中的实验结果表明,FLAM优于当前最先进的RL方法,显著提升了稳定性和整体性能。
原文摘要 · Abstract (English)
Humanoid robots have attracted significant attention in recent years. Reinforcement Learning (RL) is one of the main ways to control the whole body of humanoid robots. RL enables agents to complete tasks by learning from environment interactions, guided by task rewards. However, existing RL methods rarely explicitly consider the impact of body stability on humanoid locomotion and manipulation. Achieving high performance in whole-body control remains a challenge for RL methods that rely solely on task rewards. In this paper, we propose a Foundation model-based method for humanoid Locomotion And Manipulation (FLAM for short). FLAM integrates a stabilizing reward function with a basic policy. The stabilizing reward function is designed to encourage the robot to learn stable postures, thereby accelerating the learning process and facilitating task completion. Specifically, the robot pose is first mapped to the 3D virtual human model. Then, the human pose is stabilized and reconstructed through a human motion reconstruction model. Finally, the pose before and after reconstruction is used to compute the stabilizing reward. By combining this stabilizing reward with the task reward, FLAM effectively guides policy learning. Experimental results on a humanoid robot benchmark demonstrate that FLAM outperforms state-of-the-art RL methods, highlighting its effectiveness in improving stability and overall performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。