用自监督学习加速足式机器人地形适应运动控制训练。
BarlowWalk: Self-supervised Representation Learning for Legged Robot Terrain-adaptive Locomotion
- 引入Barlow Twins构建解耦隐空间,从历史观测中提取低维表示。
- 仅需本体感知信息即可实现连续时间自监督,减少对环境感知依赖。
- 在复杂地形下训练速度显著提升,优于先进算法的对比验证。
强化学习(RL)作为数据驱动方法,已成为解决足式机器人运动控制的有效方案。然而,当前主流的双足机器人地形穿越方法(如师生策略知识蒸馏)存在训练时间长的问题,制约开发效率。为此,本文提出BarlowWalk,一种结合自监督表示学习的改进近端策略优化(PPO)方法。该方法采用Barlow Twins算法构建解耦隐空间,将历史观测序列映射为低维表示,并实现自监督学习。同时,智能体仅需本体感知信息即可在连续时间步上完成自监督,大幅降低对外部地形感知的依赖。仿真实验表明,该方法在复杂地形场景中具有显著优势。为增强评估可信度,研究通过与先进算法的对比测试验证了所提方法的有效性。
原文摘要 · Abstract (English)
Reinforcement learning (RL), driven by data-driven methods, has become an effective solution for robot leg motion control problems. However, the mainstream RL methods for bipedal robot terrain traversal, such as teacher-student policy knowledge distillation, suffer from long training times, which limit development efficiency. To address this issue, this paper proposes BarlowWalk, an improved Proximal Policy Optimization (PPO) method integrated with self-supervised representation learning. This method employs the Barlow Twins algorithm to construct a decoupled latent space, mapping historical observation sequences into low-dimensional representations and implementing self-supervision. Meanwhile, the actor requires only proprioceptive information to achieve self-supervised learning over continuous time steps, significantly reducing the dependence on external terrain perception. Simulation experiments demonstrate that this method has significant advantages in complex terrain scenarios. To enhance the credibility of the evaluation, this study compares BarlowWalk with advanced algorithms through comparative tests, and the experimental results verify the effectiveness of the proposed method.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。