构建真实世界多粒度视觉语言定位数据集,实现零样本机器人抓取
RealVLG-R1: A Large-Scale Real-World Visual-Language Grounding Benchmark for Robotic Perception and Manipulation
- 融合多粒度标注与强化微调,统一处理语言指令下的视觉定位与抓取
- 包含165,000张图像、110亿次抓取示例,支持细粒度语义理解
- 适用于语言驱动的机器人感知与操控研究,开源可复现
视觉语言定位旨在建立自然语言与视觉实体之间的语义对应关系,使模型能根据文本指令准确识别并定位目标物体。现有方法多聚焦于粗粒度的物体级定位,而传统机器人抓取方法主要依赖几何线索,缺乏语言引导,难以应用于语言驱动的操控场景。为此,我们提出RealVLG框架,整合RealVLG-11B数据集与RealVLG-R1模型,统一真实世界视觉语言定位与抓取任务。RealVLG-11B数据集提供包括边界框、分割掩码、抓取姿态、接触点及人工验证的细粒度语言描述在内的多粒度标注,覆盖约16.5万张图像、800余种物体实例、130万条分割、检测与语言标注,以及约110亿个抓取示例。基于该数据集,RealVLG-R1在预训练大规模视觉语言模型基础上采用强化微调,给定自然语言指令时统一预测边界框、分割掩码、抓取姿态与接触点。实验表明,RealVLG可在真实未见环境中实现零样本感知与操控,建立了一个统一的语义-视觉多模态基准,为语言驱动的机器人感知与抓取策略学习提供了全面的数据与评估平台。所有数据与代码公开于https://github.com/lif314/RealVLG-R1。
原文摘要 · Abstract (English)
Visual-language grounding aims to establish semantic correspondences between natural language and visual entities, enabling models to accurately identify and localize target objects based on textual instructions. Existing VLG approaches focus on coarse-grained, object-level localization, while traditional robotic grasping methods rely predominantly on geometric cues and lack language guidance, which limits their applicability in language-driven manipulation scenarios. To address these limitations, we propose the RealVLG framework, which integrates the RealVLG-11B dataset and the RealVLG-R1 model to unify real-world visual-language grounding and grasping tasks. RealVLG-11B dataset provides multi-granularity annotations including bounding boxes, segmentation masks, grasp poses, contact points, and human-verified fine-grained language descriptions, covering approximately 165,000 images, over 800 object instances, 1.3 million segmentation, detection, and language annotations, and roughly 11 billion grasping examples. Building on this dataset, RealVLG-R1 employs Reinforcement Fine-tuning on pretrained large-scale vision-language models to predict bounding boxes, segmentation masks, grasp poses, and contact points in a unified manner given natural language instructions. Experimental results demonstrate that RealVLG supports zero-shot perception and manipulation in real-world unseen environments, establishing a unified semantic-visual multimodal benchmark that provides a comprehensive data and evaluation platform for language-driven robotic perception and grasping policy learning. All data and code are publicly available at https://github.com/lif314/RealVLG-R1.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。