arXiv:2608.12063cs.ROcs.AI2026-08

用模拟专家生成数据,让机器人用稀疏奖励快速学会又跑又抓。

Learning Loco-Manipulation From SMPC Demonstrations With Sparse Offline-to-Online RL

论文配图:Learning Loco-Manipulation From SMPC Demonstrations With Sparse Offline-to-Online RL
图 1 · 摘自论文原文
  • 用模拟模型预测控制生成海量仿真数据,自动解决探索难题。
  • 仅用稀疏任务奖励训练,学习效率显著提升,无需手动调奖励。
  • 可部署到不同机器人形态,实现超越原始控制器的最优行为。

将运动与操作结合是实现机器人自主性的关键,但标准强化学习在复杂任务上的扩展严重受限于密集奖励设计的缓慢且手动的过程。为突破这一瓶颈,我们完全在仿真中利用基于样本的模型预测控制(SMPC)作为自动化、快速可调的专家,生成大规模离线数据集。由于该数据解决了根本的探索问题,我们可仅使用稀疏的任务奖励训练离策略强化学习代理,大幅缩短新技能的学习时间,并消除对人工调参的需求。将此高层代理与低层动态稳定性控制器结合,得到的行为更优且严格符合真实任务目标,最终使学习到的策略超越原始最优控制教师。我们通过在不同形态的机器人上成功部署复杂运动-操作技能(包括带机械臂的Spot四足机器人和G1人形机器人),验证了该仿真到现实框架的鲁棒性。

原文摘要 · Abstract (English)

Integrating locomotion and manipulation is essential for robot autonomy, but scaling standard Reinforcement Learning (RL) to complex tasks is severely bottlenecked by the slow, manual process of dense reward shaping. To bypass this limitation, we leverage Sample-based Model Predictive Control (SMPC) entirely in simulation as an automated, rapidly tunable expert to generate massive offline datasets. Because this data solves the fundamental exploration problem, we can train an off-policy RL agent using purely sparse task rewards, drastically reducing the time required to learn new skills and eliminating the need for manual tuning. Integrating this high-level agent with a low-level dynamic stability controller yields more optimal behaviors that strictly align with true task objectives, ultimately allowing the learned policies to surpass the original optimal control teacher. We validate the robustness of this sim-to-real framework by successfully deploying complex loco-manipulation skills across different morphologies, including an arm-equipped Spot quadruped and a G1 humanoid.

机器人强化学习仿生控制端到端

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。