arXiv:2508.09071cs.RO2025-08被引 58

让机器人理解3D空间,用视觉+深度信息提升操作精度

GeoVLA: Empowering 3D Representations in Vision-Language-Action Models

  • 融合2D图像与3D点云,生成多模态感知表征
  • 在仿真和真实场景中实现顶尖表现,尤其擅长高度与尺度适应
  • 适合关注机器人具身智能与空间理解的研究者

视觉-语言-动作(VLA)模型为机器人执行语言指令并预测动作提供了新路径。然而,现有VLA模型主要依赖2D视觉输入,忽略了物理世界中丰富的三维几何信息,限制了其空间感知与适应能力。本文提出GeoVLA,一种新型VLA框架,通过有效整合3D信息来提升机器人操作能力。该框架使用视觉语言模型(VLM)处理图像与语言指令,提取融合的视觉-语言嵌入;同时将深度图转换为点云,并采用定制化的点编码器——点嵌入网络(Point Embedding Network),独立生成3D几何嵌入。随后,两类嵌入被拼接并输入我们提出的空间感知动作专家——3D增强型动作专家(3D-enhanced Action Expert),融合多传感器信息以生成精确动作序列。在仿真与真实环境中的大量实验表明,GeoVLA在LIBERO和ManiSkill2仿真基准上达到当前最优性能,并在需高度适应性、尺度感知与视角不变性的实际任务中展现出显著鲁棒性。

原文摘要 · Abstract (English)

Vision-Language-Action (VLA) models have emerged as a promising approach for enabling robots to follow language instructions and predict corresponding actions. However, current VLA models mainly rely on 2D visual inputs, neglecting the rich geometric information in the 3D physical world, which limits their spatial awareness and adaptability. In this paper, we present GeoVLA, a novel VLA framework that effectively integrates 3D information to advance robotic manipulation. It uses a vision-language model (VLM) to process images and language instructions,extracting fused vision-language embeddings. In parallel, it converts depth maps into point clouds and employs a customized point encoder, called Point Embedding Network, to generate 3D geometric embeddings independently. These produced embeddings are then concatenated and processed by our proposed spatial-aware action expert, called 3D-enhanced Action Expert, which combines information from different sensor modalities to produce precise action sequences. Through extensive experiments in both simulation and real-world environments, GeoVLA demonstrates superior performance and robustness. It achieves state-of-the-art results in the LIBERO and ManiSkill2 simulation benchmarks and shows remarkable robustness in real-world tasks requiring height adaptability, scale awareness and viewpoint invariance.

机器人3D理解多模态动作预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。