arXiv:2508.09547cs.CVcs.AI2025-08ACL被引 8

仅用视觉画面生成导航指令,让AI像人一样边走边想。

GoViG: Goal-Conditioned Visual Navigation Instruction Generation via Multimodal Reasoning

  • 分两步:先预测中间视觉状态,再生成连贯指令。
  • 在合成与真实场景中均超越现有方法,指标提升显著。
  • 适合做视觉导航、多模态推理的开发者和研究者。

我们提出目标条件视觉导航指令生成(GoViG),旨在仅基于初始与目标状态的自我中心视觉观测,生成语境连贯的导航指令。不同于依赖语义标注或环境地图等结构化输入的以往方法,GoViG完全利用原始自我中心视觉数据,提升了对未见和非结构化环境的适应性。该方法将任务分解为两个相互关联的子任务:(1) 导航可视化,预测连接初始与目标视图的中间视觉状态;(2) 指令生成,基于观察到和预期的视觉内容合成连贯指令。两个子任务集成于一个自回归多模态大模型中,通过定制目标确保空间准确性和语言清晰性。此外,我们引入两种多模态推理策略——单次与交替推理,模拟人类导航的认知逐步过程。为全面评估方法,我们构建了R2R-Goal数据集,融合多样化的合成与真实世界轨迹。实验结果表明,在BLEU-4和CIDEr评分上显著优于现有最先进方法,并展现出强大的跨领域泛化能力。

原文摘要 · Abstract (English)

We introduce Goal-Conditioned Visual Navigation Instruction Generation (GoViG), a new task that aims to generate contextually coherent navigation instructions solely from egocentric visual observations of initial and goal states. Unlike prior work relying on structured inputs, such as semantic annotations or environmental maps, GoViG exclusively leverages raw egocentric visual data, improving adaptability to unseen and unstructured environments. Our method addresses this task by decomposing it into two interconnected subtasks: (1) navigation visualization, predicting intermediate visual states bridging the initial and goal views; and (2) instruction generation, synthesizing coherent instructions grounded in observed and anticipated visuals. Both subtasks are integrated within an autoregressive multimodal LLM trained with tailored objectives to ensure spatial accuracy and linguistic clarity. Furthermore, we introduce two multimodal reasoning strategies, one-pass and interleaved reasoning, to mimic incremental human navigation cognition. To comprehensively evaluate our method, we propose the R2R-Goal dataset, combining diverse synthetic and real-world trajectories. Empirical results demonstrate significant performance improvements over state-of-the-art methods in BLEU-4 and CIDEr scores along with robust cross-domain generalization.

视觉导航指令生成多模态推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。