用人类手势指导机器人抓取,实现零样本泛化。
GAT-Grasp: Gesture-Driven Affordance Transfer for Task-Aware Robotic Grasping
- 通过手势与物体功能的隐含关联,从视频中提取抓取知识
- 无需预设物体先验,在新物体和杂乱环境中实现可靠抓取
- 适合需要灵活交互的协作机器人场景
实现跨多样化物体与环境的精准、可泛化抓取是智能协同机器人系统的关键。然而,现有方法常因抓取意图模糊和对未见物体适应性差,导致抓取效果不佳。本文提出GAT-Grasp,一种基于手势驱动的抓取框架,直接利用人类手部动作引导生成任务相关的抓取姿态,包括恰当的位置与朝向。具体而言,引入基于检索的属性迁移范式,利用手部动作与物体功能之间的隐含关联,从大规模人-物交互视频中提取抓取知识。通过摆脱对预设物体先验的依赖,GAT-Grasp实现了对新物体和复杂环境的零样本泛化。真实世界评估验证了其在多样且未见过场景中的鲁棒性,证明其在复杂任务设置下具备可靠的抓取能力。
原文摘要 · Abstract (English)
Achieving precise and generalizable grasping across diverse objects and environments is essential for intelligent and collaborative robotic systems. However, existing approaches often struggle with ambiguous affordance reasoning and limited adaptability to unseen objects, leading to suboptimal grasp execution. In this work, we propose GAT-Grasp, a gesture-driven grasping framework that directly utilizes human hand gestures to guide the generation of task-specific grasp poses with appropriate positioning and orientation. Specifically, we introduce a retrieval-based affordance transfer paradigm, leveraging the implicit correlation between hand gestures and object affordances to extract grasping knowledge from large-scale human-object interaction videos. By eliminating the reliance on pre-given object priors, GAT-Grasp enables zero-shot generalization to novel objects and cluttered environments. Real-world evaluations confirm its robustness across diverse and unseen scenarios, demonstrating reliable grasp execution in complex task settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。