arXiv:2503.13966cs.CVcs.RO2025-03被引 18

FlexVLN让导航模型跨数据集通用,靠大模型提升适应力。

FlexVLN: Flexible Adaptation for Diverse Vision-and-Language Navigation Tasks

  • 分层架构融合监督学习导航与大模型推理规划。
  • 在三个新数据集上表现超越以往方法,泛化能力显著提升。
  • 适合研究多场景智能体导航与大模型应用的学者。

视觉-语言导航(VLN)的目标是构建具备强适应性的具身智能体,能无缝迁移导航能力至不同任务。尽管近年进展显著,多数方法仍需针对特定数据集训练,难以在包含不同类型指令的多样化数据集间泛化。大语言模型(LLMs)展现出卓越的推理与泛化能力,在机器人动作规划中潜力巨大。本文提出FlexVLN,一种创新的分层式VLN方法,结合基于监督学习的指令跟随器的基础导航能力与大模型规划器的强泛化能力,实现跨多种VLN数据集的有效泛化。此外,设计验证机制与多模型融合机制,以缓解大模型规划器可能产生的幻觉,提升指令跟随器执行准确性。以REVERIE、SOON和CVDN-target作为域外数据集评估泛化性能,FlexVLN的表现大幅优于所有先前方法。

原文摘要 · Abstract (English)

The aspiration of the Vision-and-Language Navigation (VLN) task has long been to develop an embodied agent with robust adaptability, capable of seamlessly transferring its navigation capabilities across various tasks. Despite remarkable advancements in recent years, most methods necessitate dataset-specific training, thereby lacking the capability to generalize across diverse datasets encompassing distinct types of instructions. Large language models (LLMs) have demonstrated exceptional reasoning and generalization abilities, exhibiting immense potential in robot action planning. In this paper, we propose FlexVLN, an innovative hierarchical approach to VLN that integrates the fundamental navigation ability of a supervised-learning-based Instruction Follower with the robust generalization ability of the LLM Planner, enabling effective generalization across diverse VLN datasets. Moreover, a verification mechanism and a multi-model integration mechanism are proposed to mitigate potential hallucinations by the LLM Planner and enhance execution accuracy of the Instruction Follower. We take REVERIE, SOON, and CVDN-target as out-of-domain datasets for assessing generalization ability. The generalization performance of FlexVLN surpasses that of all the previous methods to a large extent.

视觉语言导航大模型泛化能力具身智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。