arXiv:2501.17403cs.CVcs.AI2025-01ICLR被引 18

让导航模型在真实场景中持续学习,提升长期表现。

General Scene Adaptation for Vision-and-Language Navigation

  • 设计新任务GSA-VLN,支持模型在特定场景中持续适应。
  • 构建GSA-R2R数据集,大幅提升场景与指令多样性。
  • 提出GR-DUET方法,结合记忆图与场景特训,性能领先。

视觉-语言导航(VLN)通常评估代理在多环境中的单次指令执行能力,目标是实现零样本泛化。然而现实中的导航机器人常运行于物理布局、视觉观察和语言风格相对稳定的持久环境中,现有任务设置与实际存在差距。为此,我们提出GSA-VLN新任务:要求代理在特定场景中执行指令的同时,持续适应该场景以提升长期性能。为评估此任务,需解决现有数据集中缺乏域外(OOD)数据及每场景指令数量少、风格单一的问题。因此,我们构建了新数据集GSA-R2R,显著扩展了R2R数据集的场景与指令多样性,支持在已知(ID)与域外(OOD)环境下评估适应能力。此外,设计三阶段指令重构流程,利用大模型优化原始指令,并通过角色扮演生成不同说话风格,反映用户个性化表达。我们在GSA-R2R上开展广泛实验,验证数据集有效性并对比多种方法。基于结果,提出新型方法GR-DUET,融合基于记忆的导航图与场景专属训练策略,在所有GSA-R2R划分上达到当前最优表现。

原文摘要 · Abstract (English)

Vision-and-Language Navigation (VLN) tasks mainly evaluate agents based on one-time execution of individual instructions across multiple environments, aiming to develop agents capable of functioning in any environment in a zero-shot manner. However, real-world navigation robots often operate in persistent environments with relatively consistent physical layouts, visual observations, and language styles from instructors. Such a gap in the task setting presents an opportunity to improve VLN agents by incorporating continuous adaptation to specific environments. To better reflect these real-world conditions, we introduce GSA-VLN, a novel task requiring agents to execute navigation instructions within a specific scene and simultaneously adapt to it for improved performance over time. To evaluate the proposed task, one has to address two challenges in existing VLN datasets: the lack of OOD data, and the limited number and style diversity of instructions for each scene. Therefore, we propose a new dataset, GSA-R2R, which significantly expands the diversity and quantity of environments and instructions for the R2R dataset to evaluate agent adaptability in both ID and OOD contexts. Furthermore, we design a three-stage instruction orchestration pipeline that leverages LLMs to refine speaker-generated instructions and apply role-playing techniques to rephrase instructions into different speaking styles. This is motivated by the observation that each individual user often has consistent signatures or preferences in their instructions. We conducted extensive experiments on GSA-R2R to thoroughly evaluate our dataset and benchmark various methods. Based on our findings, we propose a novel method, GR-DUET, which incorporates memory-based navigation graphs with an environment-specific training strategy, achieving state-of-the-art results on all GSA-R2R splits.

视觉语言导航场景适应大模型持续学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。