无需标注数据,用大模型生成符合任务的灵巧抓取姿势。
ZeroDexGrasp: Zero-Shot Task-Oriented Dexterous Grasp Synthesis with Prompt-Based Multi-Stage Semantic Reasoning
- 用提示词分阶段推理任务与物体语义,生成初始抓取姿态。
- 通过接触引导优化,提升抓取的物理可行性和任务匹配度。
- 零样本泛化到未见物体和复杂任务,适合智能机器人抓取场景。
任务导向的灵巧抓取在机器人操作与人机交互中具有广泛前景。然而,现有方法仍难以跨不同物体和任务指令进行泛化,因其高度依赖昂贵的标注数据以确保任务语义对齐。本文提出零样本任务导向灵巧抓取框架 ZeroDexGrasp,融合多模态大语言模型与抓取优化,生成符合特定任务目标和物体功能的人类级抓取姿态。具体而言,ZeroDexGrasp 采用提示词驱动的多阶段语义推理,从任务与物体语义中推断初始抓取配置和接触信息,再通过接触引导的抓取优化,提升姿态的物理可行性与任务一致性。实验表明,ZeroDexGrasp 能在多种未见物体类别和复杂任务需求下实现高质量的零样本灵巧抓取,推动更通用、更智能的机器人抓取发展。
原文摘要 · Abstract (English)
Task-oriented dexterous grasping holds broad application prospects in robotic manipulation and human-object interaction. However, most existing methods still struggle to generalize across diverse objects and task instructions, as they heavily rely on costly labeled data to ensure task-specific semantic alignment. In this study, we propose \textbf{ZeroDexGrasp}, a zero-shot task-oriented dexterous grasp synthesis framework integrating Multimodal Large Language Models with grasp refinement to generate human-like grasp poses that are well aligned with specific task objectives and object affordances. Specifically, ZeroDexGrasp employs prompt-based multi-stage semantic reasoning to infer initial grasp configurations and object contact information from task and object semantics, then exploits contact-guided grasp optimization to refine these poses for physical feasibility and task alignment. Experimental results demonstrate that ZeroDexGrasp enables high-quality zero-shot dexterous grasping on diverse unseen object categories and complex task requirements, advancing toward more generalizable and intelligent robotic grasping.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。