arXiv:2608.24959cs.ROcs.CV2026-08中稿 · BMVC 2026

用3D高斯表示视觉空间结构,提升机器人动作预测的几何推理能力。

GaussVLA: Geometry-Aware Spatial Reasoning for Vision-Language-Action Model

论文配图:GaussVLA: Geometry-Aware Spatial Reasoning for Vision-Language-Action Model
图 1 · 摘自论文原文
  • 将视觉与深度特征转化为紧凑的3D高斯令牌,保留几何结构信息。
  • 在LIBERO上达93.5%平均成功率,空间任务全成功,仅用200M参数。
  • 适合需要高效精准空间推理的机器人控制场景。

视觉-语言-动作(VLA)模型将视觉观测编码为无内在几何结构的2D图像块,而稠密单目深度仅注入每像素标量值,无法表达表面朝向或几何置信度,导致策略缺乏结构化空间推理能力。本文提出GaussVLA,一种基于Mamba的VLA模型,引入两个自定义模块:高斯空间分词器(GST)将冻结的语义与深度特征升维为紧凑的3D高斯令牌,并通过学习查询聚合几何显著区域;以及深度感知思维链(DA-CoT),在语言和时序流条件下进行结构化、非自回归的几何推理。在仿真与真实世界评估中,GaussVLA展现出强大的空间操作性能,同时保持参数高效。在LIBERO数据集上,平均成功率达93.5%,空间任务套件成功率为100.0%,仅使用200M参数,相比SpatialVLA相对平均成功率提升19.7%,且参数更少。

原文摘要 · Abstract (English)

Vision-Language-Action (VLA) models encode visual observations as flat 2D patch tokens that carry no intrinsic geometric structure, and augmenting them with dense monocular depth injects per-pixel scalar values that encode neither surface orientation nor geometric confidence. This leaves the policy with limited structured spatial reasoning for action prediction. We propose GaussVLA, a Mamba-based VLA that incorporates two custom modules: Gaussian Spatial Tokenizer (GST) to lift frozen semantic and depth features into compact 3D Gaussian tokens, pools geometrically salient regions with learned queries, and \emph{Depth-Aware Chain-of-Thought (DA-CoT)} that performs structured, non-autoregressive geometric reasoning under language and flow-time conditioning. Across both simulation and real-world evaluations, GaussVLA demonstrates strong spatial-manipulation performance while remaining parameter-efficient. On LIBERO, it achieves 93.5% average success and 100.0% success on the Spatial suite with only 200M parameters, improving over SpatialVLA by 19.7% relative average success while remaining significantly more parameter-efficient.

机器人控制空间推理3D高斯多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。