arXiv:2409.17621cs.RO2024-09ICRA被引 6

无需训练,用视觉语言模型结合几何信息自动选最佳递物姿势。

Leveraging Semantic and Geometric Information for Zero-Shot Robot-to-Human Handover

  • 用视觉语言模型识别物体可握区域,结合自定义提示提升定位精度。
  • 综合抓取距离与角度优化,成功率显著提升,避免干扰用户舒适区。
  • 零样本设计,适合新物品或陌生环境,服务机器人可快速部署。

人机交互涵盖多种协作任务,递物是最基础的一类。随着机器人融入人类环境,服务机器人辅助递物的潜力日益凸显。在机器人向人类递物(R2H)场景中,选择最优抓取方式至关重要,需避开人类偏好的抓握区域并减少对工作空间的侵入。现有方法或忽略几何信息,或依赖数据驱动,难以泛化至多样物体。为此,本文提出一种新颖的零样本系统,融合语义与几何信息生成最优递物抓取。首先利用视觉语言模型(VLMs)的语义知识识别抓取区域,并通过定制化视觉提示实现更精细的区域定位;再基于抓取距离与接近角度选择最优抓取点,以提升用户舒适度并避免干扰。通过消融实验与真实场景对比验证,结果表明该系统显著提升递物成功率,提供更受用户欢迎的交互体验。视频、附录等详见 https://sites.google.com/view/vlm-handover/。

原文摘要 · Abstract (English)

Human-robot interaction (HRI) encompasses a wide range of collaborative tasks, with handover being one of the most fundamental. As robots become more integrated into human environments, the potential for service robots to assist in handing objects to humans is increasingly promising. In robot-to-human (R2H) handover, selecting the optimal grasp is crucial for success, as it requires avoiding interference with the humans preferred grasp region and minimizing intrusion into their workspace. Existing methods either inadequately consider geometric information or rely on data-driven approaches, which often struggle to generalize across diverse objects. To address these limitations, we propose a novel zero-shot system that combines semantic and geometric information to generate optimal handover grasps. Our method first identifies grasp regions using semantic knowledge from vision-language models (VLMs) and, by incorporating customized visual prompts, achieves finer granularity in region grounding. A grasp is then selected based on grasp distance and approach angle to maximize human ease and avoid interference. We validate our approach through ablation studies and real-world comparison experiments. Results demonstrate that our system improves handover success rates and provides a more user-preferred interaction experience. Videos, appendixes and more are available at https://sites.google.com/view/vlm-handover/.

人机交互零样本递物机器人视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。