arXiv:2505.18418cs.RO2025-05被引 2

让四足机器人学会跨形态自适应行走,一策通用。

McARL:Morphology-Control-Aware Reinforcement Learning for Generalizable Quadrupedal Locomotion

  • 用形态向量控制策略,让模型学会适配不同体型的机器人
  • 训练后零样本迁移至不同机器人,最高速度达3.5米/秒
  • 适合需要快速部署到多种机器人的研发团队

我们提出形态感知强化学习(McARL),以解决超参数调优和迁移性能下降的问题,实现跨机器人形态的通用运动能力。通过在策略网络中引入从预定义形态范围采样的随机形态向量,使策略能学习适用于相似特征机器人的通用参数。实验表明,仅在Unitree Go1上训练的单一策略,无需重新训练或微调即可零样本迁移至其他形态(如Go2),最高实现3.5米/秒的速度;在原生Go1上可达6.0米/秒,并可泛化至A1与Mini Cheetah。相比传统PPO方法,McARL在Go2、Mini Cheetah和A1上的迁移性能提升44%-150%。

原文摘要 · Abstract (English)

We present Morphology-Control-Aware Reinforcement Learning (McARL), a new approach to overcome challenges of hyperparameter tuning and transfer loss, enabling generalizable locomotion across robot morphologies. We use a morphology-conditioned policy by incorporating a randomized morphology vector, sampled from a defined morphology range, into both the actor and critic networks. This allows the policy to learn parameters that generalize to robots with similar characteristics. We demonstrate that a single policy trained on a Unitree Go1 robot using McARL can be transferred to a different morphology (e.g., Unitree Go2 robot) and can achieve zero-shot transfer velocity of up to 3.5 m/s without retraining or fine-tuning. Moreover, it achieves 6.0 m/s on the training Go1 robot and generalizes to other morphologies like A1 and Mini Cheetah. We also analyze the impact of morphology distance on transfer performance and highlight McARL's advantages over prior approaches. McARL achieves 44-150% higher transfer performance on Go2, Mini Cheetah, and A1 compared to PPO variants.

四足机器人强化学习零样本迁移形态泛化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。