arXiv:2608.28140cs.RO2026-08中稿 · ICRA

用接触引导探索提升非抓取操作的稳定性和泛化能力

Contact-Guided Exploration for Non-Prehensile Locomanipulation with Multi-Critic RL

论文配图:Contact-Guided Exploration for Non-Prehensile Locomanipulation with Multi-Critic RL
图 1 · 摘自论文原文
  • 引入多评论家强化学习框架,通过接触奖励引导末端执行器靠近有效接触点
  • 在盒推、椅搬和洗碗机开门任务中实现稳定非抓取操作,真实机器人实验验证可行性
  • 无需特定物体设计,可通用多种几何形状,适合移动操作机器人应用

非抓取操作能灵活移动和重排重型或大体积物体,尤其结合移动操作平台时优势显著。但模型基于与模型无关的方法均面临复杂混合动力学及接触稀疏性挑战。为此,我们提出一种在多评论家强化学习框架中实现的接触引导探索策略。专门训练一个探索评论家,采用密集的接触寻求奖励,引导末端执行器向有意义的接触点靠近;其影响随训练逐步衰减,以恢复任务最优策略。候选交互点由通用抓取算法生成,使探索机制可泛化至不同物体几何形态。我们在多个任务上评估该方法,包括盒子推动、椅子搬运和洗碗机开启。最终通过四足移动操作机器人进行大量实测,验证了真实世界中可部署的非抓取操作能力。

原文摘要 · Abstract (English)

Non-prehensile manipulation offers versatile skills for moving and rearranging heavy or bulky objects, particularly when combined with a mobile manipulation platform. However, both model-based and model-free approaches struggle with the complex hybrid dynamics and the sparsity of the contact in these tasks. To address these challenges, we propose a contact-guided exploration strategy implemented within a Multi-Critic Reinforcement Learning (RL) framework. A dedicated exploration critic is trained with a dense contact-seeking reward that guides the end-effector toward meaningful contact points; its influence is progressively decayed to recover a task-optimal policy. We obtain candidate interaction points from a general-purpose grasping algorithm, enabling the exploration mechanism to generalise across various object geometries. We evaluate the approach on multiple tasks, including box pushing, chair transportation, and a dishwasher opening task. Finally, we validate the chair transportation policy through extensive experiments on a quadrupedal mobile manipulator, demonstrating deployable non-prehensile manipulation in the real world.

非抓取操作强化学习移动机器人接触引导

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。