arXiv:2607.12356cs.RO2026-07

用3D高斯点构建语义几何认知地图,提升机器人操作的三维理解能力。

VistaVLA: Geometry- and Semantic-Aware 3D Gaussian-Grounded VLA for Robotic Manipulation

论文配图:VistaVLA: Geometry- and Semantic-Aware 3D Gaussian-Grounded VLA for Robotic Manipulation
图 1 · 摘自论文原文
  • 通过3D高斯原语融合多视角视觉语言特征,生成具几何锚定的语义标记。
  • 提出MtQ机制压缩99%的高斯点,保留关键空间布局与语义信息。
  • 实测在7项任务中成功率提升22.8%,尤其在分布外任务表现突出。

视觉-语言-动作(VLA)模型通过直接将语言指令和2D视觉输入映射为动作,成为机器人操作的强大端到端范式。然而,这些模型缺乏显式的场景级3D表示,限制了对空间布局和几何约束的推理能力。尽管近期工作引入深度图或点云等显式3D线索以提升几何感知,但主要捕捉低层结构,缺乏高层语义在3D空间中的扎根。人类认知中,与物理世界的交互依赖于整合空间布局与语义上下文的3D语义认知地图——一种可持久、视角不变的内部心智模型。为此,我们提出VistaVLA,一种新颖的两阶段框架,利用3D高斯原语构建几何与语义感知的3D认知表示,并将其作为紧凑上下文标记用于VLA策略学习。具体而言,VistaVLA将多视角视觉语言特征提升至3D高斯原语,形成几何锚定的语义标记,实现视图一致的空间定位与2D视觉特征空间的对齐。为使该3D表示在有效VLA控制中计算可行,我们引入合并后查询(MtQ)机制,将密集高斯原语压缩为高度紧凑的空间信息标记,实现99%的标记减少率,同时保持动作相关的3D布局与语义上下文。在模拟与真实世界环境中的大量评估表明,该方法效果显著。值得注意的是,在真实场景中,VistaVLA在七项任务上平均成功率提升22.8%,在挑战性的分布外任务上相比VLA-Adapter基线提升30.0%。

原文摘要 · Abstract (English)

Vision-Language-Action (VLA) models have emerged as a powerful end-to-end paradigm for robotic manipulation by mapping language instructions and 2D visual inputs directly to actions. However, these models lack an explicit, scene-level 3D representation, limiting their ability to reason over spatial layouts and geometric constraints. While recent efforts incorporate explicit 3D cues, such as depth maps or point clouds, to improve geometric awareness, they primarily capture low-level structures and lack high-level semantic grounding in 3D space. In human cognition, interaction with the physical world relies on a 3D semantic cognitive map - an internal mental model that integrates spatial layouts with semantic context to enable persistent, viewpoint-invariant reasoning. In light of this, we present VistaVLA, a novel two-stage framework that constructs a geometry- and semantics-aware 3D cognitive representation from 3D Gaussian primitives and grounds it as compact context tokens for VLA policy learning. Specifically, VistaVLA lifts multi-view vision-language features into 3D Gaussian primitives, forming geometry-anchored semantic tokens that align view-consistent spatial grounding with 2D visual feature spaces. To make this 3D representation computationally tractable for effective VLA control, we introduce Merge-then-Query (MtQ), a token summarization mechanism. MtQ compresses dense Gaussian primitives into a highly compact set of spatially informative tokens, achieving a 99% token reduction while preserving action-relevant 3D layouts and semantic context. Extensive evaluations in both simulated and real-world environments demonstrate the effectiveness of VistaVLA. Notably, in real-world scenarios, VistaVLA improves success rates by 22.8% across seven real-world tasks and by 30.0% over the VLA-Adapter baseline on challenging out-of-distribution tasks.

机器人操作3D高斯视觉语言动作语义地图

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。