arXiv:2504.21530cs.ROcs.CV2025-04CVPR被引 59

用视觉语言模型生成的掩码指导机器人抓取,提升泛化能力。

RoboGround: Robotic Manipulation with Grounded Vision-Language Priors

  • 用视觉语言模型生成的掩码作为中间表示,提供目标位置与形状信息。
  • 在模拟数据上训练的策略在真实世界任务中泛化性能显著提升。
  • 适合研究机器人视觉理解与跨场景泛化的新手或进阶者。

近期机器人操作研究显示,中间表征对提升策略泛化能力具有潜力。本文探索将掩码作为有效中间表征,兼具两项优势:(1) 提供空间引导,明确目标物体和放置区域,同时传递物体形状与尺寸信息;(2) 由大规模视觉-语言模型在多样化标注数据集上预训练,具备广泛泛化能力。我们提出 RoboGround,一个基于掩码的机器人操作框架,利用掩码指导策略网络完成物体操作任务。为进一步增强泛化能力,我们设计了自动化流程,生成包含多样物体和指令的大规模模拟数据。大量实验验证了该数据集的价值及掩码作为中间引导的有效性,显著提升了机器人策略的泛化能力。

原文摘要 · Abstract (English)

Recent advancements in robotic manipulation have highlighted the potential of intermediate representations for improving policy generalization. In this work, we explore grounding masks as an effective intermediate representation, balancing two key advantages: (1) effective spatial guidance that specifies target objects and placement areas while also conveying information about object shape and size, and (2) broad generalization potential driven by large-scale vision-language models pretrained on diverse grounding datasets. We introduce RoboGround, a grounding-aware robotic manipulation system that leverages grounding masks as an intermediate representation to guide policy networks in object manipulation tasks. To further explore and enhance generalization, we propose an automated pipeline for generating large-scale, simulated data with a diverse set of objects and instructions. Extensive experiments show the value of our dataset and the effectiveness of grounding masks as intermediate guidance, significantly enhancing the generalization abilities of robot policies.

机器人操作视觉语言掩码引导

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。