arXiv:2605.24642cs.CVcs.RO2026-05被引 2

探究几何模型如何提升视觉语言动作模型的3D理解能力

Understanding the Impact of Geometric Foundation Models on Vision-Language-Action Models

论文配图:Understanding the Impact of Geometric Foundation Models on Vision-Language-Action Models
图 1 · 摘自论文原文
  • 用线性探测量化现有模型的几何理解短板
  • 对比三种几何注入架构,发现融合方式影响显著
  • 揭示训练数据与重建质量对性能的关键作用

近期研究探索了视觉语言动作模型(VLAs)与几何基础模型(GFMs)在3D重建中的结合潜力,如VGGT。尽管融合后的几何VLAs性能有所提升,但仍有三个关键问题未明:(i) 现有VLAs是否已具备足够的几何理解能力?(ii) 如何最优地将几何理解注入VLAs?(iii) 其他设计选择(如训练数据、相机数量、重建质量)如何影响性能?本文以GR00T-N1.5为VLA、VGGT为GFM,进行严谨实验分析。首先,通过线性探测首次量化了VLAs与GFMs之间的“几何差距”。其次,实现三种不同几何注入架构,在保持底层细节一致的前提下,公平比较其效果。最后,系统分析非架构因素对几何VLAs性能的影响。

原文摘要 · Abstract (English)

Recent work explores new opportunities at the intersection of vision-language-action models (VLAs) and geometric foundation models (GFMs) for 3D reconstruction, such as VGGT. While the resulting geometric VLAs often show improved performance, it remains unclear (i) if modern VLAs already have sufficient geometric understanding to start with, (ii) what is the best architecture to inject geometric understanding into a VLA, and (iii) what is the effect of other design choices that affect geometric VLAs. In this paper we provide a rigorous experimental analysis to shed light on these questions, for a specific choice of VLA (GR00T-N1.5) and GFM (VGGT). Our first contribution is to formalize prior work's intuition that current VLAs lack geometric understanding, by providing a rigorous analysis based on linear probing. The analysis quantifies, for the first time, the "geometric gap" between VLAs and GFMs. Our second contribution is to identify and compare different strategies to bridge GFMs with VLAs. We implement three different architectures, which differ in the way they inject geometry in the VLA, while keeping low-level implementation details as similar as possible, to ensure a fair comparison. Finally, we analyze the impact of non-architectural choices (e.g., training data, number of cameras, reconstruction quality) on the performance of the geometric VLAs.

几何建模视觉语言动作基础模型3D重建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。