arXiv:2505.08574cs.RO2025-05被引 2

用专家控制数据训练单一神经网络,实现四足机器人多步态实时控制。

End-to-End Multi-Task Policy Learning from NMPC for Quadruped Locomotion

  • 基于NMPC专家示范,端到端学习多任务策略。
  • 在仿真与真实硬件上均实现高精度步态复现和无缝切换。
  • 统一策略简化部署,适合复杂地形实时运动控制。

四足机器人在复杂非结构化环境中表现优异,但其非线性动力学、高自由度及实时控制的计算需求使其高效自适应运动仍具挑战。基于优化的控制器如非线性模型预测控制(NMPC)性能良好,但依赖精确状态估计且计算开销大,难以在真实场景部署。本文提出一种多任务学习(MTL)框架,利用专家级NMPC示范,训练单一神经网络直接从原始本体感觉输入预测多种运动行为的动作。我们在四足机器人Go1上进行了仿真与真实硬件的全面评估,结果表明该方法能准确复现专家行为,实现平滑步态切换,并简化实时控制流程。所提出的MTL架构使统一策略中学习多种步态成为可能,在所有任务中均达到高$R^{2}$分数(预测关节目标的拟合优度)。

原文摘要 · Abstract (English)

Quadruped robots excel in traversing complex, unstructured environments where wheeled robots often fail. However, enabling efficient and adaptable locomotion remains challenging due to the quadrupeds' nonlinear dynamics, high degrees of freedom, and the computational demands of real-time control. Optimization-based controllers, such as Nonlinear Model Predictive Control (NMPC), have shown strong performance, but their reliance on accurate state estimation and high computational overhead makes deployment in real-world settings challenging. In this work, we present a Multi-Task Learning (MTL) framework in which expert NMPC demonstrations are used to train a single neural network to predict actions for multiple locomotion behaviors directly from raw proprioceptive sensor inputs. We evaluate our approach extensively on the quadruped robot Go1, both in simulation and on real hardware, demonstrating that it accurately reproduces expert behavior, allows smooth gait switching, and simplifies the control pipeline for real-time deployment. Our MTL architecture enables learning diverse gaits within a unified policy, achieving high $R^{2}$ scores for predicted joint targets across all tasks.

四足机器人多任务学习神经控制端到端

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。