arXiv:2511.17411cs.ROcs.LG2025-11被引 2

用3D标注的非机器人图像数据提升视觉语言模型,让机器人更懂空间。

SPEAR-1: Scaling Beyond Robot Demonstrations via 3D Understanding

  • 用2D图像加3D标注训练3D感知模型,无需大量机器人数据
  • 仅用20分之一的机器人示范,性能超过或媲美顶尖模型
  • 适合研究通用机器人控制和少样本学习的研究者

机器人基础模型(RFM)有望成为通用、端到端的机器人控制系统,但其在新环境、任务和机械体上的泛化能力仍受限。我们指出,主要瓶颈在于基础:多数RFM基于互联网预训练的视觉-语言模型(VLM),而这些模型在2D图像-语言任务上训练,缺乏实现3D世界具身控制所需的3D空间推理能力。直接用大规模机器人数据填补这一差距成本高且难扩展。因此,我们提出用易获取的非机器人图像数据加入3D标注,增强预训练VLM的3D理解能力。基于此策略,我们训练了SPEAR-VLM——一个能从单张2D图像推断物体3D坐标的3D感知VLM。在此基础上,我们提出主要贡献SPEAR-1:集成3D感知与语言指令控制的机器人基础模型。该模型在约4500万帧、来自24个Open X-Embodiment数据集的视频上训练,性能优于或媲美π₀-FAST和π₀.₅等先进模型,同时仅需20倍更少的机器人示范。这一精心设计的训练策略解锁了新的VLM能力,显著提升了具身控制的可靠性,超越仅依赖机器人数据的局限。模型权重与3D标注数据集已公开于https://spear.insait.ai。

原文摘要 · Abstract (English)

Robotic Foundation Models (RFMs) hold great promise as generalist, end-to-end systems for robot control. Yet their ability to generalize across new environments, tasks, and embodiments remains limited. We argue that a major bottleneck lies in their foundations: most RFMs are built by fine-tuning internet-pretrained Vision-Language Models (VLMs). However, these VLMs are trained on 2D image-language tasks and lack the 3D spatial reasoning inherently required for embodied control in the 3D world. Bridging this gap directly with large-scale robotic data is costly and difficult to scale. Instead, we propose to enrich easy-to-collect non-robotic image data with 3D annotations and enhance a pretrained VLM with 3D understanding capabilities. Following this strategy, we train SPEAR-VLM, a 3D-aware VLM that infers object coordinates in 3D space from a single 2D image. Building on SPEAR-VLM, we introduce our main contribution, $~\textbf{SPEAR-1}$: a robotic foundation model that integrates grounded 3D perception with language-instructed embodied control. Trained on $\sim$45M frames from 24 Open X-Embodiment datasets, SPEAR-1 outperforms or matches state-of-the-art models such as $π_0$-FAST and $π_{0.5}$, while it uses 20$\times$ fewer robot demonstrations. This carefully-engineered training strategy unlocks new VLM capabilities and as a consequence boosts the reliability of embodied control beyond what is achievable with only robotic data. We make our model weights and 3D-annotated datasets publicly available at https://spear.insait.ai.

机器人控制3D理解视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。