无需参考模型,一策适配多形态四足机器人运动控制
Reference Free Platform Adaptive Locomotion for Quadrupedal Robots using a Dynamics Conditioned Policy
- 用动态条件化策略让单一控制策略适配不同体型和动力学特性
- 在ANYmal C上实现零样本迁移,速度跟踪误差降低30%
- 适合需跨平台通用控制的机器人研发与部署场景
本文提出平台自适应运动控制(PAL),一种适用于不同形态和动力学特性的四足机器人统一控制方法。通过深度强化学习,在程序生成的多种机器人上训练单一运动策略,该策略基于本体感知状态与基底速度指令,输出期望关节动作目标,并利用局部系统动态的隐式嵌入进行条件化。探索了两种条件化策略:基于GRU的动态编码器与基于形态属性估计器,实验表明形态感知条件化在速度任务跟踪上优于时间动态编码。两种方法均在多个未见过的仿真四足机器人上实现稳健的零样本迁移。进一步表明,训练中引入多样化的机器人形态与动力学可显著提升泛化能力,使速度跟踪误差相比基线方法降低最高达30%。尽管PAL未在所有情况下超越最优无参考控制器,但其分析揭示了关键设计选择,为现有技术改进提供了指导。
原文摘要 · Abstract (English)
This article presents Platform Adaptive Locomotion (PAL), a unified control method for quadrupedal robots with different morphologies and dynamics. We leverage deep reinforcement learning to train a single locomotion policy on procedurally generated robots. The policy maps proprioceptive robot state information and base velocity commands into desired joint actuation targets, which are conditioned using a latent embedding of the temporally local system dynamics. We explore two conditioning strategies - one using a GRU-based dynamics encoder and another using a morphology-based property estimator - and show that morphology-aware conditioning outperforms temporal dynamics encoding regarding velocity task tracking for our hardware test on ANYmal C. Our results demonstrate that both approaches achieve robust zero-shot transfer across multiple unseen simulated quadrupeds. Furthermore, we demonstrate the need for careful robot reference modelling during training: exposing the policy to a diverse set of robot morphologies and dynamics leads to improved generalization, reducing the velocity tracking error by up to 30% compared to the baseline method. Despite PAL not surpassing the best-performing reference-free controller in all cases, our analysis uncovers critical design choices and informs improvements to the state of the art.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。