arXiv:2512.18938cs.RO2025-12

开源框架实现四足机械臂强化学习控制的仿真到现实迁移

A Framework for Deploying Learning-based Quadruped Loco-Manipulation

  • 基于ROS构建统一管道,打通Sim-to-Sim与Sim-to-Real全流程
  • 在Isaac Gym和MuJoCo间发现接触模型差异影响策略表现
  • 实机试验证明全身协同控制提升抓取范围与操作精度

四足移动操作机器人具备敏捷运动与操作潜力,但其控制与从仿真到现实的迁移仍具挑战。强化学习(RL)有望实现全身控制,但多数框架为专有且难以在真实硬件上复现。本文提出一个开源管道,用于在Unitree B1四足机器人搭载Z1机械臂上训练、评测和部署基于RL的控制器。该框架通过ROS统一仿真到仿真(sim-to-sim)与仿真到现实(sim-to-real)迁移路径,重实现了Isaac Gym中的策略,并通过硬件抽象层扩展至MuJoCo,最终部署于物理硬件。仿真对比显示Isaac Gym与MuJoCo的接触模型存在差异,影响策略行为;真实世界遥控抓取实验表明,全身协同控制显著扩展了操作范围并优于浮基基线方法。该管道提供透明可复现的开发分析基础,将开源以支持后续研究。

原文摘要 · Abstract (English)

Quadruped mobile manipulators offer strong potential for agile loco-manipulation but remain difficult to control and transfer reliably from simulation to reality. Reinforcement learning (RL) shows promise for whole-body control, yet most frameworks are proprietary and hard to reproduce on real hardware. We present an open pipeline for training, benchmarking, and deploying RL-based controllers on the Unitree B1 quadruped with a Z1 arm. The framework unifies sim-to-sim and sim-to-real transfer through ROS, re-implementing a policy trained in Isaac Gym, extending it to MuJoCo via a hardware abstraction layer, and deploying the same controller on physical hardware. Sim-to-sim experiments expose discrepancies between Isaac Gym and MuJoCo contact models that influence policy behavior, while real-world teleoperated object-picking trials show that coordinated whole-body control extends reach and improves manipulation over floating-base baselines. The pipeline provides a transparent, reproducible foundation for developing and analyzing RL-based loco-manipulation controllers and will be released open source to support future research.

强化学习四足机器人仿真迁移机器人控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。