arXiv:2502.08054cs.ROcs.LG2025-02被引 3

用双臂协作解决被遮挡物体抓取难题,提升机器人在复杂环境中的操作成功率。

COMBO-Grasp: Learning Constraint-Based Manipulation for Bimanual Occluded Grasping

  • 双策略协同:约束策略自监督学习稳定姿态,抓取策略强化学习优化动作
  • 价值函数引导协调,使双臂配合更高效,成功率显著提升
  • 适合需要双臂操作的工业场景,尤其适用于真实世界部署

本文针对遮挡环境下机器人抓取难题,即因环境约束导致目标抓取位姿无法实现的情况。传统方法难以处理非握持或双臂协作等人类常用策略,而现有强化学习方法因任务复杂性不适用,示范学习则需大量专家数据且难获取。为此,受人类双臂协作启发,提出基于约束的双臂遮挡抓取方法COMBO-Grasp,采用两个协同策略:利用自监督数据训练的约束策略生成稳定姿态,以及通过强化学习训练的抓取策略完成重定向与抓取。关键创新在于价值函数引导的策略协调机制,在抓取策略训练中通过联合训练的价值函数梯度优化约束策略输出,增强双臂协同与任务表现。最后,采用教师-学生策略蒸馏,将基于点云的策略有效部署至真实环境。实验表明,COMBO-Grasp在模拟和真实环境中均显著优于对比基线,并成功泛化到未见物体。

原文摘要 · Abstract (English)

This paper addresses the challenge of occluded robot grasping, i.e. grasping in situations where the desired grasp poses are kinematically infeasible due to environmental constraints such as surface collisions. Traditional robot manipulation approaches struggle with the complexity of non-prehensile or bimanual strategies commonly used by humans in these circumstances. State-of-the-art reinforcement learning (RL) methods are unsuitable due to the inherent complexity of the task. In contrast, learning from demonstration requires collecting a significant number of expert demonstrations, which is often infeasible. Instead, inspired by human bimanual manipulation strategies, where two hands coordinate to stabilise and reorient objects, we focus on a bimanual robotic setup to tackle this challenge. In particular, we introduce Constraint-based Manipulation for Bimanual Occluded Grasping (COMBO-Grasp), a learning-based approach which leverages two coordinated policies: a constraint policy trained using self-supervised datasets to generate stabilising poses and a grasping policy trained using RL that reorients and grasps the target object. A key contribution lies in value function-guided policy coordination. Specifically, during RL training for the grasping policy, the constraint policy's output is refined through gradients from a jointly trained value function, improving bimanual coordination and task performance. Lastly, COMBO-Grasp employs teacher-student policy distillation to effectively deploy point cloud-based policies in real-world environments. Empirical evaluations demonstrate that COMBO-Grasp significantly improves task success rates compared to competitive baseline approaches, with successful generalisation to unseen objects in both simulated and real-world environments.

双臂抓取强化学习机器人操作点云感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。