arXiv:2512.13080cs.RO2025-12被引 12

让机器人模型学会从2D视频理解3D空间,提升动作准确性

Spatial-Aware VLA Pretraining through Visual-Physical Alignment from Human Videos

  • 用人类示范视频提取3D视觉与动作标注,对齐视觉与物理空间
  • 在下游任务中显著提升2D视觉与3D动作的对齐效果
  • 适合做具身智能、机器人视觉-语言-动作学习的研究者

视觉-语言-动作(VLA)模型通过整合视觉感知与语言引导策略学习,为机器人学习提供了新范式。然而,现有方法大多依赖2D视觉输入在3D物理环境中执行动作,导致感知与动作之间的语义脱节。为此,我们提出一种空间感知的VLA预训练范式,在预训练阶段显式对齐视觉空间与物理空间,使模型在学习机器人策略前就具备3D空间理解能力。基于预训练的视觉-语言模型,我们利用大规模人类示范视频提取3D视觉与3D动作标注,构建一种新的监督信号,实现2D视觉观测与3D空间推理的对齐。我们以VIPA-VLA为例,采用双编码器架构,引入3D视觉编码器,增强语义视觉表示的3D感知特征。在下游机器人任务中,VIPA-VLA显著提升了2D视觉与3D动作的对齐性,生成更鲁棒、泛化性更强的机器人策略。

原文摘要 · Abstract (English)

Vision-Language-Action (VLA) models provide a promising paradigm for robot learning by integrating visual perception with language-guided policy learning. However, most existing approaches rely on 2D visual inputs to perform actions in 3D physical environments, creating a significant gap between perception and action grounding. To bridge this gap, we propose a Spatial-Aware VLA Pretraining paradigm that performs explicit alignment between visual space and physical space during pretraining, enabling models to acquire 3D spatial understanding before robot policy learning. Starting from pretrained vision-language models, we leverage large-scale human demonstration videos to extract 3D visual and 3D action annotations, forming a new source of supervision that aligns 2D visual observations with 3D spatial reasoning. We instantiate this paradigm with VIPA-VLA, a dual-encoder architecture that incorporates a 3D visual encoder to augment semantic visual representations with 3D-aware features. When adapted to downstream robot tasks, VIPA-VLA achieves significantly improved grounding between 2D vision and 3D action, resulting in more robust and generalizable robotic policies.

机器人学习空间对齐多模态预训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。