arXiv:2604.12509cs.ROcs.CV2026-04被引 1

用离线强化学习优化低质量控制器,实现机器人全身协同操作

Whole-Body Mobile Manipulation using Offline Reinforcement Learning on Sub-optimal Controllers

  • 先用随机化轻量级控制器生成多样演示数据
  • 离线强化学习提升策略,在仿真中成功率超基准方法
  • 无需真人操作数据,直接部署到真实机器人上

移动操作(MoMa)如开门、拉抽屉等任务需机器人基座与机械臂的全身协同。传统全身体控制器(WBC)依赖繁琐调参且易失效;学习方法虽泛化能力强,但通常需昂贵的全程遥操作数据或复杂奖励设计。我们发现,即使次优的WBC也具备强结构先验:可在任务相关状态-动作空间内高效收集数据,并通过离线强化学习进一步优化。基于此,提出WHOLE-MoMa两阶段框架:首先随机化轻量级WBC生成多样化示范,再利用离线强化学习通过奖励信号识别并拼接更优行为。为支持复杂协调任务所需的动作分块扩散策略,扩展了离线隐式Q学习,引入分块级评论器评估和优势加权策略提取。在模拟的TIAGo++机器人上,三个逐步增加难度的任务中,该方法显著优于WBC、行为克隆及多个离线强化学习基线。策略可直接迁移至真实机器人,无需微调,在双臂抽屉操作中达80%成功率,在同步柜门开启与物体放置任务中达68%成功率,全程未使用任何遥操作或真实世界训练数据。

原文摘要 · Abstract (English)

Mobile Manipulation (MoMa) of articulated objects, such as opening doors, drawers, and cupboards, demands simultaneous, whole-body coordination between a robot's base and arms. Classical whole-body controllers (WBCs) can solve such problems via hierarchical optimization, but require extensive hand-tuned optimization and remain brittle. Learning-based methods, on the other hand, show strong generalization capabilities but typically rely on expensive whole-body teleoperation data or heavy reward engineering. We observe that even a sub-optimal WBC is a powerful structural prior: it can be used to collect data in a constrained, task-relevant region of the state-action space, and its behavior can still be improved upon using offline reinforcement learning. Building on this, we propose WHOLE-MoMa, a two-stage pipeline that first generates diverse demonstrations by randomizing a lightweight WBC, and then applies offline RL to identify and stitch together improved behaviors via a reward signal. To support the expressive action-chunked diffusion policies needed for complex coordination tasks, we extend offline implicit Q-learning with Q-chunking for chunk-level critic evaluation and advantage-weighted policy extraction. On three tasks of increasing difficulty using a TIAGo++ mobile manipulator in simulation, WHOLE-MoMa significantly outperforms WBC, behavior cloning, and several offline RL baselines. Policies transfer directly to the real robot without finetuning, achieving 80% success in bimanual drawer manipulation and 68% in simultaneous cupboard opening and object placement, all without any teleoperated or real-world training data.

机器人操作离线强化学习全身控制仿真到现实

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。