arXiv:2605.28548cs.CV2026-05被引 4

用生成深度图提升机器人视觉语言模型的实操能力

GEM: Generative Supervision Helps Embodied Intelligence

论文配图:GEM: Generative Supervision Helps Embodied Intelligence
图 1 · 摘自论文原文
  • 在预训练中加入深度图生成任务,融合空间物理信息
  • 新数据集GEM-4M含400万条带深度监督的多模态数据
  • 在仿真与真实场景中均显著提升任务执行效果

具身视觉语言模型在机器人领域表现出色,尤其在视觉-语言-动作框架中。然而,标准文本引导预训练侧重高层语义,难以满足具身环境中的低层空间与物理知识需求。本文提出GEM,一种生成式监督的具身视觉语言模型,将深度图生成任务直接融入预训练阶段。通过联合训练该生成目标,显著提升了模型的语义理解与物理操作能力。为此,我们构建并发布了GEM-4M——一个大规模综合数据集,包含400万条融合定位、推理与规划任务的数据,并配有高质量深度监督。大量实验表明,GEM在多个具身基准上达到领先水平。此外,部署的动作模型GEM-VLA在仿真与真实世界评估中均展现出卓越的任务执行能力。代码、模型与数据集详见https://zhaorw02.github.io/GEM/

原文摘要 · Abstract (English)

Embodied Vision-Language Models (VLMs) have demonstrated impressive performance and generalization in robotics, particularly within Vision-Language-Action frameworks. However, a significant gap remains between the high-level semantic focus of standard text-guided pre-training paradigms and the low-level spatial and physical knowledge critical for execution in embodied environments. In this paper, we introduce GEM, a Generative-supervised Embodied vision-language Model designed to bridge this divide. We propose integrating a depth map generation task directly into the VLM pre-training phase. By training this generative objective jointly with the main model, we observe substantial improvements in embodied intelligence, significantly enhancing both semantic understanding and physical operation capabilities. To support this paradigm, we curate and release GEM-4M, a comprehensive large-scale dataset featuring a mixture of grounding, reasoning, and planning data paired with high-quality depth supervision. Extensive experiments demonstrate that GEM achieves state-of-the-art results across diverse embodied benchmarks. Furthermore, our deployed action model, GEM-VLA, exhibits vastly superior task execution abilities in both simulation environments and real-world evaluations. Code, models, and datasets are available at https://zhaorw02.github.io/GEM/

具身智能视觉语言深度生成机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。