arXiv:2603.26412cs.RO2026-03中稿 · Robotics and Auton…被引 1

用大模型构建物体部件知识库,实现跨任务泛化抓取

Generalizable task-oriented object grasping through LLM-guided ontology and similarity-based planning

  • 用大语言模型构建人机共理解的物体功能部件知识体系
  • 通过几何特征匹配实现对新物体部件的精准识别与抓取规划
  • 无需依赖视觉语义,适配新物体和复杂任务场景

任务导向型抓取(TOG)比简单抓取更具挑战性,需精确识别物体部件并选择合适的抓取位置以确保有效操作。现有方法虽结合视觉-语言模型实现部件级分割与任务感知抓取规划,但其在部件识别和抓取推断上存在不稳定性,难以泛化至多样物体与任务。为此,我们提出一种不依赖视觉语义特征的几何中心策略,克服传统模型对视角敏感的问题。核心包括:1)基于大语言模型构建支持直观人类指令的功能部件-任务本体;2)采用采样式几何分析方法,结合多点分布与距离度量从点云中定位选定部件;3)设计相似性匹配框架,利用已有分割与抓取知识的相似物体作为参考,指导未知目标的抓取规划。实验证明该方法在功能部件选择、识别与抓取生成方面具有高精度。同时,通过扩展本体知识,成功实现对新类别物体的泛化,展现出对多种物体与任务的良好适应性。

原文摘要 · Abstract (English)

Task-oriented grasping (TOG) is more challenging than simple object grasping because it requires precise identification of object parts and careful selection of grasping areas to ensure effective and robust manipulation. While recent approaches have trained large-scale vision-language models to integrate part-level object segmentation with task-aware grasp planning, their instability in part recognition and grasp inference limits their ability to generalize across diverse objects and tasks. To address this issue, we introduce a novel, geometry-centric strategy for more generalizable TOG that does not rely on semantic features from visual recognition, effectively overcoming the viewpoint sensitivity of model-based approaches. Our main proposals include: 1) an object-part-task ontology for functional part selection based on intuitive human commands, constructed using a Large Language Model (LLM); 2) a sampling-based geometric analysis method for identifying the selected object part from observed point clouds, incorporating multiple point distribution and distance metrics; and 3) a similarity matching framework for imitative grasp planning, utilizing similar known objects with pre-existing segmentation and grasping knowledge as references to guide the planning for unknown targets. We validate the high accuracy of our approach in functional part selection, identification, and grasp generation through real-world experiments. Additionally, we demonstrate the method's generalization capabilities to novel-category objects by extending existing ontological knowledge, showcasing its adaptability to a broad range of objects and tasks.

抓取规划大模型本体几何分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。