arXiv:2512.21970cs.RO2025-12被引 13

用立体视觉提升机器人动作模型的空间感知能力

StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision

  • 设计双路视觉编码器,同时提取几何与语义信息
  • 真实场景下任务成功率提升33.4%,抗视角变化能力强
  • 适合需要精准空间判断的机器人操控场景

尽管视觉-语言-动作(VLA)模型在通用操作任务中表现优异,但普遍存在细粒度空间感知不足和视角鲁棒性差的问题。这主要源于依赖预训练的RGB编码器,其缺乏显式几何信息,更注重语义对齐而非几何表征。本文提出StereoVLA,首个引入大规模合成立体数据以获取丰富几何线索的VLA模型。该模型采用几何-语义联合编码器(GeoSem),通过微小视差差异提取精确空间信息,同时保留像素级语义特征,支持语言驱动的操作。此外,设计两种协同训练目标:交互区域深度估计用于精准空间推理,相机参数估计用于隐式对齐感知与动作坐标系。在真实世界实验中,相比使用多种输入模态的基线模型,StereoVLA成功率达33.4%绝对提升,并展现出对近半球视角的鲁棒性。

原文摘要 · Abstract (English)

While Vision-Language-Action (VLA) models excel in generalist manipulation, they often lack fine-grained spatial awareness and show limited viewpoint robustness. This limitation largely stems from the reliance on pretrained RGB encoders, which lack explicit geometric cues and prioritize semantic alignment over geometric representation. We argue that effective visual representations for VLA models must jointly encode both semantic and geometric information. In this paper, we introduce StereoVLA, the first VLA model to incorporate rich geometric cues from large-scale synthetic stereo data. StereoVLA employs a Geometric-and-Semantic (GeoSem) vision encoder that extracts geometric cues from subtle stereo-view disparities for precise spatial perception, while simultaneously capturing semantic features from pixel observations to support language-conditioned manipulation. Additionally, we introduce two synergistic co-training objectives: Interaction-Region Depth Estimation for precise spatial reasoning, and Camera Parameter Estimation to implicitly align perception and action coordinate systems. Compared with baselines that employ various input modalities, StereoVLA achieves a 33.4% absolute gain in success rate in real-world experiments and demonstrates robustness to near-hemispheric camera perspectives. Project page: https://shengliangd.github.io/StereoVLA-Webpage.

机器人操控立体视觉空间感知多模态模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。