arXiv:2607.05122cs.CVcs.RO2026-07中稿 · RSS 2026 workshop

用分割图标注可走/不可走区域,让机器人导航更准

Green for Go, Red for No: Visual Grounding via Semantic Segmentation for VLA Navigation Policies

论文配图:Green for Go, Red for No: Visual Grounding via Semantic Segmentation for VLA Navigation Policies
图 1 · 摘自论文原文
  • 用SegFormer实时分割场景,绿标可走区,红标不可走区
  • 长指令下路径误差降27%-44%,路径长度减少30%
  • 适合提升长指令导航,无需重训练模型

视觉-语言-动作(VLA)模型能根据自然语言和视觉目标实现机器人导航,但易受感知干扰和场景理解模糊影响。本文首次对VLA导航策略的视觉接地进行实证评估。提出一种基于SegFormer的实时分割接地方法,以绿色标记可通行区域,红色标记不可通行区域。评估了两种变体:仅观察分割与观察-目标联合增强。在Grand Tour数据集上使用OmniVLA模型,结果显示,在最远航点处,视觉接地使平均航点误差降低27%-44%,且长指令下收益更大,图像目标任务中改善有限。归一化误差分析表明,接地主要作为轨迹长度正则器,使预测路径长度减少30%,但未提升单位距离推理能力。结果表明,视觉接地是一种简单、低开销的改进方式,无需模型重训,但无法弥补分布外指令中缺失的训练信号。

原文摘要 · Abstract (English)

Vision-language-action (VLA) models enable robot navigation from natural language and visual goals, but remain susceptible to perceptual distractions and ambiguous scene interpretations. This paper presents the first empirical evaluation of visual grounding for VLA navigation policies. We propose a real-time segmentation-based grounding method that highlights traversable areas in green and non-traversable areas in red using SegFormer. Two variants are evaluated: observation-only segmentation and joint observation-goal augmentation. Using OmniVLA on the Grand Tour dataset, we show that visual grounding reduces the mean waypoint error by 27-44% at the farthest waypoint, depending on the instruction length. The benefits are greater for long instructions than for short instructions, and grounding provides little improvement for image goals. Normalized error analysis indicates that grounding primarily acts as a trajectory length regularizer, reducing the predicted path length by 30% without improving per-unit-distance reasoning. Our results indicate that visual grounding offers a simple, computationally inexpensive method to improve VLA navigation without model retraining, although it cannot compensate for missing training signals in out-of-distribution instructions.

视觉接地机器人导航分割模型多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。