用视觉语言模型解决复杂遮挡下的6自由度抓取问题
VISO-Grasp: Vision-Language Informed Spatial Object-centric 6-DoF Active View Planning and Grasping in Clutter and Invisibility
- 融合视觉语言模型进行空间推理与主动视角规划
- 实测抓取成功率87.5%,尝试次数最少
- 适合机器人在遮挡或完全不可见环境下抓取
我们提出VISO-Grasp,一种新型的视觉-语言感知系统,用于系统性应对严重遮挡环境中的可见性约束。通过利用基础模型(FMs)进行空间推理与主动视角规划,该框架构建并持续更新实例中心的空间关系表示,提升了在复杂遮挡下的抓取成功率。此外,该表示支持主动最优视角(NBV)规划,并在直接抓取不可行时优化序列抓取策略。我们还引入多视角不确定性驱动的抓取融合机制,实时提升抓取置信度与方向不确定性估计,确保抓取执行的鲁棒性与稳定性。大量真实世界实验表明,VISO-Grasp在目标导向抓取中达到87.5%的成功率,且所需抓取尝试次数最少,显著优于基线方法。据我们所知,VISO-Grasp是首个将基础模型统一整合到目标感知主动视角规划与6-DoF抓取中的框架,适用于存在严重遮挡和完全不可见性约束的环境。代码已开源:https://github.com/YitianShi/vMF-Contact
原文摘要 · Abstract (English)
We propose VISO-Grasp, a novel vision-language-informed system designed to systematically address visibility constraints for grasping in severely occluded environments. By leveraging Foundation Models (FMs) for spatial reasoning and active view planning, our framework constructs and updates an instance-centric representation of spatial relationships, enhancing grasp success under challenging occlusions. Furthermore, this representation facilitates active Next-Best-View (NBV) planning and optimizes sequential grasping strategies when direct grasping is infeasible. Additionally, we introduce a multi-view uncertainty-driven grasp fusion mechanism that refines grasp confidence and directional uncertainty in real-time, ensuring robust and stable grasp execution. Extensive real-world experiments demonstrate that VISO-Grasp achieves a success rate of $87.5\%$ in target-oriented grasping with the fewest grasp attempts outperforming baselines. To the best of our knowledge, VISO-Grasp is the first unified framework integrating FMs into target-aware active view planning and 6-DoF grasping in environments with severe occlusions and entire invisibility constraints. Code is available at: https://github.com/YitianShi/vMF-Contact
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。