arXiv:2503.02387cs.ROcs.SY2025-03被引 1

仅用单张彩色图像实现未知物体抓取,通过超二次曲面建模提升抓取成功率。

RGBSQGrasp: Inferring Local Superquadric Primitives from Single RGB Image for Graspability-Aware Bin Picking

  • 基于超二次曲面和单目图像推断物体几何形状,无需深度传感器。
  • 在真实场景中实现92%的抓取成功率,对未知物体有良好适应性。
  • 适合机器人抓取、自动化分拣等实际应用场景。

箱内抓取因遮挡和物理限制导致视觉信息不足,难以识别与抓取物体。现有方法多依赖已知CAD模型或先验几何结构,泛化能力受限;另一些方法虽可直接从RGB-D数据回归抓取位姿,但深度传感噪声和缺乏对象理解使抓取生成与评估困难。超二次曲面(SQ)能以紧凑、可解释的方式表征物体的物理特性和可抓取性。然而,从有限视角恢复其形状仍具挑战,因传统方法需多视角点云重建,不适用于箱内抓取。为此,我们提出RGBSQGrasp:一个基于单目RGB图像的抓取框架,利用超二次曲面形状原语与基础度量深度估计模型,无需深度传感器即可推断抓取位姿。该框架包含通用跨平台数据集生成流程、基于基础模型的点云估计模块、全局-局部超二次曲面拟合网络及SQ引导的抓取位姿采样模块。通过集成各组件,RGBSQGrasp实现基于几何推理的稳定抓取,提升对未见物体的适应性。真实机器人实验表明,在密集箱体环境中达到92%抓取成功率,验证了其有效性。

原文摘要 · Abstract (English)

Bin picking is a challenging robotic task due to occlusions and physical constraints that limit visual information for object recognition and grasping. Existing approaches often rely on known CAD models or prior object geometries, restricting generalization to novel or unknown objects. Other methods directly regress grasp poses from RGB-D data without object priors, but the inherent noise in depth sensing and the lack of object understanding make grasp synthesis and evaluation more difficult. Superquadrics (SQ) offer a compact, interpretable shape representation that captures the physical and graspability understanding of objects. However, recovering them from limited viewpoints is challenging, as existing methods rely on multiple perspectives for near-complete point cloud reconstruction, limiting their effectiveness in bin-picking. To address these challenges, we propose \textbf{RGBSQGrasp}, a grasping framework that leverages superquadric shape primitives and foundation metric depth estimation models to infer grasp poses from a monocular RGB camera -- eliminating the need for depth sensors. Our framework integrates a universal, cross-platform dataset generation pipeline, a foundation model-based object point cloud estimation module, a global-local superquadric fitting network, and an SQ-guided grasp pose sampling module. By integrating these components, RGBSQGrasp reliably infers grasp poses through geometric reasoning, enhancing grasp stability and adaptability to unseen objects. Real-world robotic experiments demonstrate a 92% grasp success rate, highlighting the effectiveness of RGBSQGrasp in packed bin-picking environments.

机器人抓取单目视觉超二次曲面箱内分拣

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。