arXiv:2508.02917cs.CVcs.AI2025-08中稿 · ICNSLP 2025被引 1

用现成大模型走路线指令,发现全景动作更优但仍有差距

Following Route Instructions using Large Vision-Language Models: A Comparison between Low-level and Panoramic Action Spaces

  • 用开源大模型直接微调,不改架构也不仿真训练
  • 在R2R数据集上达到41%成功率,全景动作优于低级动作
  • 适合想快速尝试视觉语言导航的开发者参考

视觉-语言导航(VLN)旨在让自主机器人通过自然语言指令在陌生环境中导航。尽管近期大型视觉语言模型(LVLM)在此任务上展现出潜力,但大多数现有系统依赖为导航专门设计和优化的模型,导致现成LVLM的潜力未被充分挖掘。此外,早期方法使用以自身视角为基础的低级动作空间(如“左转”或“前进”),而新模型更倾向采用离散可导航视点的全景动作空间。本文研究了两个问题:(1) 现成的LVLM(未经架构修改或模拟器训练微调)能否有效支持VLN任务;(2) 这类模型是否能同时支持低级与全景动作范式。为此,我们在Room-to-Room(R2R)数据集上对开源模型Qwen2.5-VL-3B-Instruct进行微调,并在两种动作空间下评估其表现。最佳模型在R2R测试集上达到41%的成功率,表明现成的LVLM虽可学习执行视觉-语言导航,但仍落后于专为该任务设计的模型。

原文摘要 · Abstract (English)

Vision-and-Language Navigation (VLN) refers to the task of enabling autonomous robots to navigate unfamiliar environments by following natural language instructions. While recent Large Vision-Language Models (LVLMs) have shown promise in this task, most current VLM systems rely on models specifically designed and optimized for navigation, leaving the potential of off-the-shelf LVLMs underexplored. Furthermore, while older VLN approaches used low-level action spaces with egocentric views and atomic actions (such as "turn left" or "move forward"), newer models tend to favor panoramic action spaces with discrete navigable viewpoints. This paper investigates (1) whether off-the-shelf LVLMs (fine-tuned without architectural modifications or simulator-based training) can effectively support VLN tasks and (2) whether such models can support both low-level and panoramic action paradigms. To this end, we fine-tune the open-source model Qwen2.5-VL-3B-Instruct on the Room-to-Room (R2R) dataset and evaluate its empirical performance across both low-level and panoramic action spaces. The best resulting model achieves a 41% success rate on the R2R test set, demonstrating that while off-the-shelf LVLMs can learn to perform Vision-and-Language Navigation, they still lag behind models specifically designed for this task.

视觉导航大模型应用路径规划

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。