研究发现3D感知精度提升对视觉语言导航帮助有限,关键在用对核心信息。
Exploring Bottlenecks in VLM-LLM Navigation: How 3D Scene Understanding Capability Impacts Zero-Shot VLN

- 用上下文成功率上限分析导航系统瓶颈
- 感知精度超阈值后导航效果不再提升
- 应聚焦导航所需的关键空间信息
零样本视觉语言导航因数据成本低且具备泛化能力而备受关注。该任务通常依赖预训练视觉语言模型(VLM)与大语言模型(LLM)的结合:VLM构建3D场景图,LLM负责高层推理决策。然而当前存在关键瓶颈——3D感知模型过度追求像素级精度,与具身导航所需的实时性及计算限制直接冲突。本文量化了3D场景理解能力对导航性能的影响,基于典型VLM-LLM框架,为两个核心子系统建立统计成功率(SR)上界:1)依赖拓扑映射语义的慢速LLM规划器;2)利用空间坐标和边界框执行决策的快速反应导航器。使用前沿3D场景理解模型进行评估,验证了所提上界,并发现感知饱和现象:当感知精度超过某一阈值后,导航成功率提升趋于平缓。研究建议,3D场景理解应从严格像素精度转向更关注导航相关的核心词汇与准确的边界框比例。
原文摘要 · Abstract (English)
Zero-shot vision-and-language navigation (VLN) has gained significant attention due to its minimal data collection costs and inherent generalization. This paradigm is typically driven by the integration of pre-trained Vision-Language Models (VLMs) and Large Language Models (LLMs), where VLMs construct 3D scene graphs while LLMs handle high-level reasoning and decision-making. However, a critical bottleneck exists in this system: current 3D perception models prioritize pixel-level accuracy, directly conflicting with the strict computational limits and real-time efficiency demanded by embodied navigation. To address this gap, this paper quantifies the actual impact of 3D scene understanding capability on VLN performance. Based on typical VLM-LLM frameworks, we propose statistical success rate (SR) upper bounds for two core subsystems: 1) the slow LLM planner, which relies on topological mapping semantics, and 2) the fast reactive navigator, which utilizes spatial coordinates and bounding boxes to execute LLM decisions. Evaluations using state-of-the-art 3D scene understanding models validate our proposed bounds and reveal a perception saturation phenomenon, indicating that improvements in perception accuracy beyond a certain threshold yield diminishing returns in navigation success. Our findings suggest that 3D scene understanding for VLN should pivot away from strict pixel-level precision, prioritizing instead navigation-relevant core vocabularies and accurate bounding box proportions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。