用语音控制机器人灵巧抓取,复杂场景下准确率高
Grasp What You Want: Embodied Dexterous Grasping System Driven by Your Voice
- 结合语音与视觉的语义对齐方法,提升指令理解精度
- 基于人手交互原理设计抓取策略,成功率高且稳定
- 适合需要语音控制的智能机器人开发人员参考
近年来,随着机器人技术的发展,人机协作日益重要。然而,现有机器人仅依靠语音命令难以准确理解人类意图。传统夹持器和吸盘系统在非结构化环境中交互自然性差,缺乏高级操作能力且适应性不足。本文提出具身灵巧抓取系统(EDGS),旨在解决复杂场景下的物体抓取问题。我们采用视觉-语言模型(VLM)融合语音指令与视觉信息,实现目标物体多维属性的精准对齐。受人手-物体交互启发,设计了包含拇指轴线、多指环绕和指尖接触力学等原则的稳健、精确、高效的抓取策略。通过实验评估指代表达表示增强(RERE)在指代表达分割中的表现,验证了系统能准确识别并匹配指代表达。大量实验证明,EDGS可有效完成复杂抓取任务,具备高稳定性与成功率,展现出在具身智能领域的应用潜力。
原文摘要 · Abstract (English)
In recent years, as robotics has advanced, human-robot collaboration has gained increasing importance. However, current robots struggle to fully and accurately interpret human intentions from voice commands alone. Traditional gripper and suction systems often fail to interact naturally with humans, lack advanced manipulation capabilities, and are not adaptable to diverse tasks, especially in unstructured environments. This paper introduces the Embodied Dexterous Grasping System (EDGS), designed to tackle object grasping in cluttered environments for human-robot interaction. We propose a novel approach to semantic-object alignment using a Vision-Language Model (VLM) that fuses voice commands and visual information, significantly enhancing the alignment of multi-dimensional attributes of target objects in complex scenarios. Inspired by human hand-object interactions, we develop a robust, precise, and efficient grasping strategy, incorporating principles like the thumb-object axis, multi-finger wrapping, and fingertip interaction with an object's contact mechanics. We also design experiments to assess Referring Expression Representation Enrichment (RERE) in referring expression segmentation, demonstrating that our system accurately detects and matches referring expressions. Extensive experiments confirm that EDGS can effectively handle complex grasping tasks, achieving stability and high success rates, highlighting its potential for further development in the field of Embodied AI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。