arXiv:2507.00416cs.ROcs.CV2025-07被引 63

让视觉语言模型从2D图像中隐式理解3D空间,无需额外传感器

Evo-0: Vision-Language-Action Model with Implicit Spatial Understanding

  • 用现成的视觉几何模型隐式引入3D结构信息
  • 在模拟和真实场景中显著提升现有视觉语言动作模型表现
  • 无需深度传感器或预训练深度模型,适合通用机器人应用

视觉-语言-动作(VLA)模型为具备感知、推理与行动能力的通用机器人提供了有前景的框架。这些模型通常基于预训练的视觉语言模型(VLM),后者因大规模图文预训练而擅长语义理解,但普遍缺乏精确的空间理解能力,因其主要在二维图像-文本对上训练,缺乏三维监督。现有方法虽引入点云或深度图等显式3D输入,但需额外深度传感器或预训练深度估计模型,可能导致效果不佳。本文提出一种即插即用模块,通过利用现成的视觉几何基础模型,隐式将3D几何特征融入VLA模型,使其仅凭RGB图像即可获得具备深度感知的视觉表征,增强对场景几何结构及物体间空间关系的理解。我们在模拟与真实世界的一系列空间挑战任务中评估该方法,实验表明其显著提升了主流VLA模型在多样化场景中的性能。

原文摘要 · Abstract (English)

Vision-Language-Action (VLA) models have emerged as a promising framework for enabling generalist robots capable of perceiving, reasoning, and acting in the real world. These models usually build upon pretrained Vision-Language Models (VLMs), which excel at semantic understanding due to large-scale image and text pretraining. However, existing VLMs typically lack precise spatial understanding capabilities, as they are primarily tuned on 2D image-text pairs without 3D supervision. To address this limitation, recent approaches have incorporated explicit 3D inputs such as point clouds or depth maps, but this necessitates additional depth sensors or pre-trained depth estimation models, which may yield defective results. In contrast, our work introduces a plug-and-play module that implicitly incorporates 3D geometry features into VLA models by leveraging an off-the-shelf visual geometry foundation model. This integration provides the model with depth-aware visual representations, improving its ability to understand the geometric structure of the scene and the spatial relationships among objects from RGB images alone. We evaluate our method on a set of spatially challenging tasks in both simulation and the real world. Extensive evaluations show that our method significantly improves the performance of state-of-the-art VLA models across diverse scenarios.

视觉语言空间理解机器人3D感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。