用视觉引导提升机器人抓取精度与泛化能力
RoboGrasp: A Universal Grasping Policy for Robust Robotic Control
- 融合预训练抓取检测模型与扩散方法,利用视觉任务增强指导
- 少样本学习和盒形提示任务中成功率提升34%
- 适合需要高鲁棒性的复杂场景抓取应用
模仿学习与世界模型在推动可泛化机器人学习方面展现出巨大潜力,而机器人抓取仍是实现精准操作的关键挑战。现有方法通常高度依赖机械臂状态数据和RGB图像,易对特定物体形状或位置过拟合。为解决这一问题,我们提出RoboGrasp——一种集成预训练抓取检测模型的通用抓取策略框架。通过利用目标检测与分割任务提供的稳健视觉引导,RoboGrasp显著提升了抓取精度、稳定性和泛化能力,在少样本学习和盒形提示任务中成功率最高提升34%。该框架基于扩散方法构建,可适配多种机器人学习范式,实现在多样复杂场景下的精确可靠操作。此方案为应对真实世界抓取挑战提供了可扩展且多功能的解决方案。
原文摘要 · Abstract (English)
Imitation learning and world models have shown significant promise in advancing generalizable robotic learning, with robotic grasping remaining a critical challenge for achieving precise manipulation. Existing methods often rely heavily on robot arm state data and RGB images, leading to overfitting to specific object shapes or positions. To address these limitations, we propose RoboGrasp, a universal grasping policy framework that integrates pretrained grasp detection models with robotic learning. By leveraging robust visual guidance from object detection and segmentation tasks, RoboGrasp significantly enhances grasp precision, stability, and generalizability, achieving up to 34% higher success rates in few-shot learning and grasping box prompt tasks. Built on diffusion-based methods, RoboGrasp is adaptable to various robotic learning paradigms, enabling precise and reliable manipulation across diverse and complex scenarios. This framework represents a scalable and versatile solution for tackling real-world challenges in robotic grasping.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。