用离线数据训练通用四足机器人行走策略,不依赖具体机型。
GRoQ-LoCO: Generalist and Robot-agnostic Quadruped Locomotion Control using Offline Datasets
- 基于注意力机制的统一框架,从多机器人离线数据中学习通用行走策略。
- 零样本迁移至12kg Go1和70kg Stoch 5机器人,在多种地形成功行走。
- 仅使用本体感知数据,无需微调即可在i7 NUC上实时部署。
近期大规模离线训练展示了通用策略学习在复杂机器人任务中的潜力。然而,由于连续动力学及对多样地形与机器人形态的实时适应需求,该方法在腿式行走任务中仍面临挑战。本文提出GRoQ-LoCO,一种可扩展的注意力框架,仅依赖离线数据,学习单一通用四足行走策略,适用于多种机器人与地形。该方法利用来自多个四足机器人上采集的专家示范数据,涵盖楼梯穿越(非周期步态)和平地行走(周期步态),训练出支持行为融合的通用模型。关键在于,框架仅使用各机器人本体感知数据,不引入任何机器人特异性编码。策略可在Intel i7 NUC上直接部署,实现低延迟控制输出,无需测试时优化。大量实验表明,该方法实现了跨多样化四足机器人的零样本迁移,包括在商用12kg Unitree Go1上的硬件部署。特别地,在不同机器人间步态技能分布不均的设定下,仍成功将平地行走与楼梯穿越能力迁移至所有机器人。初步结果还显示,无需微调即可在70kg Stoch 5上实现平地与户外地形步行。这些成果证明了离线数据驱动学习在跨多种四足形态与行为间的泛化潜力。
原文摘要 · Abstract (English)
Recent advancements in large-scale offline training have demonstrated the potential of generalist policy learning for complex robotic tasks. However, applying these principles to legged locomotion remains a challenge due to continuous dynamics and the need for real-time adaptation across diverse terrains and robot morphologies. In this work, we propose GRoQ-LoCO, a scalable, attention-based framework that learns a single generalist locomotion policy across multiple quadruped robots and terrains, relying solely on offline datasets. Our approach leverages expert demonstrations from two distinct locomotion behaviors - stair traversal (non-periodic gaits) and flat terrain traversal (periodic gaits) - collected across multiple quadruped robots, to train a generalist model that enables behavior fusion. Crucially, our framework operates solely on proprioceptive data from all robots without incorporating any robot-specific encodings. The policy is directly deployable on an Intel i7 nuc, producing low-latency control outputs without any test-time optimization. Our extensive experiments demonstrate zero-shot transfer across highly diverse quadruped robots and terrains, including hardware deployment on the Unitree Go1, a commercially available 12kg robot. Notably, we evaluate challenging cross-robot training setups where different locomotion skills are unevenly distributed across robots, yet observe successful transfer of both flat walking and stair traversal behaviors to all robots at test time. We also show preliminary walking on Stoch 5, a 70kg quadruped, on flat and outdoor terrains without requiring any fine tuning. These results demonstrate the potential of offline, data-driven learning to generalize locomotion across diverse quadruped morphologies and behaviors.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。