arXiv:2511.18845cs.AI2025-11AAAI被引 3

让视觉语言导航同时思考画面和指令,提升复杂场景下的行走准确率。

UNeMo: Collaborative Visual-Language Reasoning and Navigation via a Multimodal World Model

  • 构建多模态世界模型,融合图像、语言和动作预测下一步视觉状态。
  • 在未见场景中导航准确率提升2.1%(R2R)和0.7%(REVERIE)。
  • 适合研究视觉语言推理与智能体导航的开发者和研究员。

视觉-语言导航(VLN)要求智能体根据视觉图像和自然语言指令自主导航复杂环境,仍具高度挑战性。近期利用预训练大语言模型(LLMs)增强语言引导导航推理的研究展现出良好前景,但其推理局限于语言模态,缺乏视觉推理能力。此外,现有推理模块与导航策略优化分离,导致目标不一致甚至冲突。为此,我们提出UNeMo框架,实现视觉状态推理与导航决策的协同优化。该框架引入多模态世界模型(MWM),以视觉特征、语言指令和导航动作为输入,联合预测后续视觉状态,支持跨模态推理。通过分层预测-反馈机制(HPN),MWM与导航策略协作:第一层基于当前视觉-语言特征生成动作;随后MWM推断动作后的视觉状态,指导第二层精细化决策。二者形成动态双向促进机制——MWM推理优化导航策略,策略决策反哺提升MWM推理精度。在R2R与REVERIE数据集上的实验表明,UNeMo在未见场景中的导航准确率分别优于当前最优方法2.1%和0.7%,验证了其有效性。

原文摘要 · Abstract (English)

Vision-and-Language Navigation (VLN) requires agents to autonomously navigate complex environments via visual images and natural language instructions--remains highly challenging. Recent research on enhancing language-guided navigation reasoning using pre-trained large language models (LLMs) has shown promising prospects. However, the reasoning of such methods is limited to the linguistic modality, lacking visual reasoning capabilities. Moreover, existing reasoning modules are optimized separately from navigation policies, leading to incompatibility and potential conflicts in optimization objectives.To tackle these challenges, we introduce UNeMo, a novel framework designed for the collaborative optimization of visual state reasoning and navigational decision-making. It introduces a Multimodal World Model (MWM) that takes visual features, language instructions, and navigational actions as inputs to jointly predict subsequent visual states, enabling cross-modal reasoning. Via a Hierarchical Prediction-Feedback (HPN) mechanism, MWM collaborates with navigation policies: the first layer generates actions using current vision-and-language features; MWM then infers post-action visual states to guide the second layer's fine-grained decisions. This forms a dynamic bidirectional promotion mechanism where MWM reasoning optimizes navigation policies, while policy decisions feedback to improve MWM's reasoning accuracy. Experiments on R2R and REVERIE datasets show UNeMo outperforms state-of-the-art methods by 2.1% and 0.7% in navigation accuracy for unseen scenes, validating its effectiveness.

视觉导航多模态世界模型语言推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。