无需训练,用视觉语言模型实现任务导向抓取
Training-free Task-oriented Grasp Generation
- 结合预训练抓取模型与视觉语言模型,实现无训练任务导向抓取
- 相比基线方法,抓取成功率提升36.9%,任务符合率显著提高
- 适合研究机器人操作与人机交互的学者快速部署应用
本文提出一种无需训练的任务导向抓取生成流程,通过融合预训练抓取模型与视觉语言模型(VLMs),在不依赖特定任务数据训练的前提下,利用VLMs的语义推理能力融入任务需求。相较于仅关注稳定抓取的传统方法,该方法在五个不同查询策略下评估,基于候选抓取的不同视觉表征进行优化,实验显示整体抓取成功率绝对提升最高达36.9%,任务符合率也显著改善。结果表明,视觉语言模型在增强任务导向操作方面具有巨大潜力,为机器人抓取与人机协作研究提供了新思路。
原文摘要 · Abstract (English)
This paper presents a training-free pipeline for task-oriented grasp generation that combines pre-trained grasp generation models with vision-language models (VLMs). Unlike traditional approaches that focus solely on stable grasps, our method incorporates task-specific requirements by leveraging the semantic reasoning capabilities of VLMs. We evaluate five querying strategies, each utilizing different visual representations of candidate grasps, and demonstrate significant improvements over a baseline method in both grasp success and task compliance rates, with absolute gains of up to 36.9\% in overall success rate. Our results underline the potential of VLMs to enhance task-oriented manipulation, providing insights for future research in robotic grasping and human-robot interaction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。