arXiv:2505.06832cs.RO2025-05被引 3

统一框架实现双臂精准抓取,支持开放词汇语义理解与几何约束。

UniDiffGrasp: A Unified Framework Integrating VLM Reasoning and VLM-Guided Part Diffusion for Open-Vocabulary Constrained Grasping with Dual Arms

  • 融合视觉语言模型与部件引导扩散,实现语义目标到抓取动作的直接映射。
  • 单臂抓取成功率0.876,双臂达0.767,显著优于现有方法。
  • 适用于复杂场景下开放词汇、带约束的双臂协同抓取任务。

开放词汇、任务导向的特定功能部件抓取,尤其在双臂协作场景中仍具挑战性。当前视觉-语言模型虽增强任务理解,却常难以生成满足约束的精确抓取姿态并实现有效双臂协调。本文提出UniDiffGrasp,一种统一框架,将视觉语言模型推理与部件引导扩散相结合。该框架利用视觉语言模型解析用户指令,识别语义目标(物体、部件、操作模式),并通过开放词汇分割实现目标定位。关键在于,识别出的部件直接为约束抓取扩散场(CGDF)提供几何约束,通过部件引导扩散实现无需重训练的高效高质量6-DoF抓取。针对双臂任务,框架定义不同目标区域,分别对每臂应用部件引导扩散,并选取稳定协同抓取策略。在真实世界部署中,单臂抓取成功率0.876,双臂达0.767,显著超越现有最先进方法,验证了其在复杂现实场景中实现精准、协调的开放词汇抓取的能力。

原文摘要 · Abstract (English)

Open-vocabulary, task-oriented grasping of specific functional parts, particularly with dual arms, remains a key challenge, as current Vision-Language Models (VLMs), while enhancing task understanding, often struggle with precise grasp generation within defined constraints and effective dual-arm coordination. We innovatively propose UniDiffGrasp, a unified framework integrating VLM reasoning with guided part diffusion to address these limitations. UniDiffGrasp leverages a VLM to interpret user input and identify semantic targets (object, part(s), mode), which are then grounded via open-vocabulary segmentation. Critically, the identified parts directly provide geometric constraints for a Constrained Grasp Diffusion Field (CGDF) using its Part-Guided Diffusion, enabling efficient, high-quality 6-DoF grasps without retraining. For dual-arm tasks, UniDiffGrasp defines distinct target regions, applies part-guided diffusion per arm, and selects stable cooperative grasps. Through extensive real-world deployment, UniDiffGrasp achieves grasp success rates of 0.876 in single-arm and 0.767 in dual-arm scenarios, significantly surpassing existing state-of-the-art methods, demonstrating its capability to enable precise and coordinated open-vocabulary grasping in complex real-world scenarios.

双臂抓取视觉语言模型扩散模型开放词汇

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。