让四足机器人带机械臂实现多任务自主操控
MLM: Learning Multi-task Loco-Manipulation Whole-Body Control for Quadruped Robot with Arm
- 结合真实与仿真数据,用强化学习训练多任务协同控制
- 通过轨迹库和课程采样机制提升学习效率与任务平衡
- 支持零样本迁移,适合需要灵活操作的机器人研发
带六自由度机械臂的四足机器人全身体运动操控仍是难题,尤其在多任务控制方面。本文提出MLM框架,结合真实世界与仿真数据的强化学习方法,使机器人能自主或在人机遥控下完成多种任务的全身运动操控。为解决多任务学习中的平衡问题,引入带有自适应课程采样机制的轨迹库,高效利用真实采集的轨迹进行学习。针对仅有历史观测的部署场景及不同空间范围任务的性能提升需求,提出轨迹-速度预测策略网络,可预测不可观测的未来轨迹与速度。借助大量仿真数据与课程化奖励设计,控制器在仿真中实现全身行为,并实现零样本迁移到真实世界。仿真消融实验验证了方法必要性与有效性,真实世界实验在配备Airbot机械臂的Go2机器人上展示了多任务执行的良好表现。
原文摘要 · Abstract (English)
Whole-body loco-manipulation for quadruped robots with arms remains a challenging problem, particularly in achieving multi-task control. To address this, we propose MLM, a reinforcement learning framework driven by both real-world and simulation data. It enables a six-DoF robotic arm-equipped quadruped robot to perform whole-body loco-manipulation for multiple tasks autonomously or under human teleoperation. To address the problem of balancing multiple tasks during the learning of loco-manipulation, we introduce a trajectory library with an adaptive, curriculum-based sampling mechanism. This approach allows the policy to efficiently leverage real-world collected trajectories for learning multi-task loco-manipulation. To address deployment scenarios with only historical observations and to enhance the performance of policy execution across tasks with different spatial ranges, we propose a Trajectory-Velocity Prediction policy network. It predicts unobservable future trajectories and velocities. By leveraging extensive simulation data and curriculum-based rewards, our controller achieves whole-body behaviors in simulation and zero-shot transfer to real-world deployment. Ablation studies in simulation verify the necessity and effectiveness of our approach, while real-world experiments on a Go2 robot with an Airbot robotic arm demonstrate the policy's good performance in multi-task execution.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。