arXiv:2409.11518cs.ROcs.CV2024-09ICRA被引 6

用轻量语义分割让机器人听懂指令并精准操作物体。

Robot Manipulation in Salient Vision through Referring Image Segmentation and Geometric Constraints

  • 将语言描述转为视觉边界精确的图像分割结果。
  • 6.6MB小模型在46项真实任务中表现超越传统方法。
  • 适合需要自然语言控制的智能机器人研发人员。

本文通过在机器人感知模块中集成紧凑型指代图像分割模型,实现基于语言上下文的真实世界机器人操作。首先提出CLIPU²Net,一种面向细粒度边界与结构分割的轻量级指代图像分割模型。随后将该模型部署于眼动伺服系统中,实现真实世界的机器人控制。系统核心在于将显著视觉信息表示为几何约束,使视觉感知直接转化为可执行命令。在46个真实世界机器人操作任务上的实验表明,该方法优于依赖人工特征标注的传统视觉伺服方法,在细粒度指代图像分割上表现优异,且解码器仅需6.6 MB存储空间,支持跨多种场景的机器人控制。

原文摘要 · Abstract (English)

In this paper, we perform robot manipulation activities in real-world environments with language contexts by integrating a compact referring image segmentation model into the robot's perception module. First, we propose CLIPU$^2$Net, a lightweight referring image segmentation model designed for fine-grain boundary and structure segmentation from language expressions. Then, we deploy the model in an eye-in-hand visual servoing system to enact robot control in the real world. The key to our system is the representation of salient visual information as geometric constraints, linking the robot's visual perception to actionable commands. Experimental results on 46 real-world robot manipulation tasks demonstrate that our method outperforms traditional visual servoing methods relying on labor-intensive feature annotations, excels in fine-grain referring image segmentation with a compact decoder size of 6.6 MB, and supports robot control across diverse contexts.

机器人操控语义分割视觉伺服

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。