分层强化学习+模型预测控制,让机器人精准触碰物体并真实世界直接可用。
Where to Touch, How to Contact: A Hierarchical RL-MPC Framework for Geometry-Aware Sim-to-Real Manipulation
- 高层用RL决定触点和目标姿态,低层用MPC实时规划接触动作。
- 比传统方法少40倍决策步数,2倍控制步数,零样本跨仿真与现实迁移。
- 适合需要精确接触的复杂操作任务,如翻转、推移、重定位物体。
在接触密集型灵巧操作中,需同时考虑全局几何与非光滑接触动力学。端到端策略虽简化问题,但依赖大量数据且仿真到现实迁移差。本文提出一种分层强化学习-模型预测控制框架:高层强化学习(RL)策略预测接触意图,即以对象为中心的界面,指定(i)物体表面接触位置和(ii)接触后的目标子姿态;基于此意图,低层接触隐式模型预测控制(MPC)优化局部接触模式,并实时重规划接触动力学,生成鲁棒的机器人动作以推动物体至子目标。我们在非抓取任务上评估,包括多种物体形状下的几何泛化推移、基于转动/翻转的物体姿态重定向、以及环境辅助的物体再定位。结果表明,该框架在显著减少数据(相比T型推移任务,减少40倍RL决策步、2倍控制步)的前提下,实现高成功率、强鲁棒性,并实现零样本仿真到现实迁移。
原文摘要 · Abstract (English)
A key challenge in contact-rich dexterous manipulation is the need to jointly reason over global geometry and nonsmooth contact dynamics. End-to-end policies bypass this complexity, but often require large amounts of data and transfer poorly from simulation to reality. We address the limitations with a simple insight: dexterous manipulation is inherently hierarchical--at a high level, a robot decides where to touch (geometry); at a low level it determines how to move the object through contact dynamics. Building on this insight, we propose a hierarchical RL--MPC framework in which a high-level reinforcement learning (RL) policy predicts a contact intention, a novel object-centric interface that specifies (i) an object-surface contact location and (ii) a post-contact object subgoal pose. Conditioned on the contact intention, a low-level contact-implicit model predictive control (MPC) optimizes local contact modes and real-time (re)plans through contact dynamics to generate robot actions that robustly move the object toward each subgoal. We evaluate the framework on non-prehensile tasks, including geometry-generalized pushing across diverse object shapes, pivoting/flipping-based object reorientation, and environment-assisted object repositioning. It achieves high success rate with substantially reduced data (40 times fewer RL decision steps and 2 times fewer control steps in the T-pushing comparison), highly robust performance, and zero-shot sim-to-real transfer.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。