arXiv:2510.12276cs.RO2025-10被引 114

让视觉语言模型学会空间感知,无需深度图也能精准执行机器人指令。

Spatial Forcing: Implicit Spatial Representation Alignment for Vision-language-action Model

  • 通过隐式对齐3D基础模型的几何表示,提升模型空间理解能力。
  • 在模拟和真实场景中超越现有2D/3D方法,动作精度显著提高。
  • 训练速度提升3.8倍,适合追求高效高精度的机器人应用研究者。

视觉-语言-动作(VLA)模型在使机器人遵循语言指令并执行精确操作方面展现出巨大潜力。然而,大多数VLA基于仅在2D数据上预训练的视觉-语言模型,缺乏准确的空间感知能力,难以在三维物理世界中有效运作。现有方法尝试引入显式的3D传感器输入(如深度图或点云),但受限于传感器噪声、硬件异构性以及数据集中的深度覆盖不全。依赖2D图像估计3D信息的方法也因深度估计算法性能有限而受限。本文提出空间强制(Spatial Forcing, SF),一种简单而有效的对齐策略,无需显式3D输入或深度估计算法,即可隐式引导VLA模型发展空间理解能力。SF将VLA的中间视觉表征与预训练3D基础模型生成的几何表示进行对齐。通过在中间层施加对齐约束,SF促使模型编码更丰富的空间信息,从而提升动作精度。大量仿真与真实环境实验表明,SF达到当前最佳性能,优于基于2D和3D的VLA方法。同时,SF将训练速度提升高达3.8倍,并在多种机器人任务中显著提升数据效率。项目页面见 https://spatial-forcing.github.io/

原文摘要 · Abstract (English)

Vision-language-action (VLA) models have recently shown strong potential in enabling robots to follow language instructions and execute precise actions. However, most VLAs are built upon vision-language models pretrained solely on 2D data, which lack accurate spatial awareness and hinder their ability to operate in the 3D physical world. Existing solutions attempt to incorporate explicit 3D sensor inputs such as depth maps or point clouds, but these approaches face challenges due to sensor noise, hardware heterogeneity, and incomplete depth coverage in existing datasets. Alternative methods that estimate 3D cues from 2D images also suffer from the limited performance of depth estimators. We propose Spatial Forcing (SF), a simple yet effective alignment strategy that implicitly forces VLA models to develop spatial comprehension capabilities without relying on explicit 3D inputs or depth estimators. SF aligns intermediate visual embeddings of VLAs with geometric representations produced by pretrained 3D foundation models. By enforcing alignment at intermediate layers, SF guides VLAs to encode richer spatial representations that enhance action precision. Extensive experiments in simulation and real-world environments demonstrate that SF achieves state-of-the-art results, surpassing both 2D- and 3D-based VLAs. SF further accelerates training by up to 3.8x and improves data efficiency across diverse robotic tasks. Project page is at https://spatial-forcing.github.io/

空间感知机器人多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。