用语言分析物体部件,生成更灵活的机器人抓取动作。
PartDexTOG: Generating Dexterous Task-Oriented Grasping via Language-driven Part Analysis
- 通过大模型分析物体部件功能,生成任务导向的抓取描述。
- 在OakInk-shape数据集上,抓取穿透体积、位移和一致性指标分别提升3.58%、2.87%、41.43%。
- 适合需要灵活抓取新物体与新任务的机器人系统应用。
任务导向抓取是机器人操作中的关键挑战。尽管已有进展,但多数方法未针对灵巧手设计。灵巧手具备更高精度与适应性,能更有效完成任务导向抓取。本文提出PartDexTOG,通过语言驱动的部件分析增强灵巧抓取。输入一个3D物体和语言描述的任务,首先由大模型生成类别级与部件级的抓取描述;随后,基于这些描述,使用类别-部件条件扩散模型为每个部件分别生成灵巧抓取动作;最后,通过几何一致性度量筛选最合理的抓取-部件组合。实验表明,该方法得益于大模型对物体部件的开放世界知识推理,在不同几何形状与任务下表现优异。在OakInk-shape数据集上优于所有现有方法,穿透体积、抓取位移和P-FID指标分别提升3.58%、2.87%、41.43%。尤其在处理新类别和新任务时展现出良好泛化能力。
原文摘要 · Abstract (English)
Task-oriented grasping is a crucial yet challenging task in robotic manipulation. Despite the recent progress, few existing methods address task-oriented grasping with dexterous hands. Dexterous hands provide better precision and versatility, enabling robots to perform task-oriented grasping more effectively. In this paper, we argue that part analysis can enhance dexterous grasping by providing detailed information about the object's functionality. We propose PartDexTOG, a method that generates dexterous task-oriented grasps via language-driven part analysis. Taking a 3D object and a manipulation task represented by language as input, the method first generates the category-level and part-level grasp descriptions w.r.t the manipulation task by LLMs. Then, a category-part conditional diffusion model is developed to generate a dexterous grasp for each part, respectively, based on the generated descriptions. To select the most plausible combination of grasp and corresponding part from the generated ones, we propose a measure of geometric consistency between grasp and part. We show that our method greatly benefits from the open-world knowledge reasoning on object parts by LLMs, which naturally facilitates the learning of grasp generation on objects with different geometry and for different manipulation tasks. Our method ranks top on the OakInk-shape dataset over all previous methods, improving the Penetration Volume, the Grasp Displace, and the P-FID over the state-of-the-art by $3.58\%$, $2.87\%$, and $41.43\%$, respectively. Notably, it demonstrates good generality in handling novel categories and tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。