arXiv:2504.16516cs.CVcs.AI2025-04被引 6

分层融合多模态信息,让导航机器人更懂指令。

Think Hierarchically, Act Dynamically: Hierarchical Multi-modal Fusion and Reasoning for Vision-and-Language Navigation

  • 分层次融合视觉、语言和历史信息,捕捉跨模态复杂交互
  • 在REVERIE、R2R、SOON数据集上表现优于现有方法
  • 适合研究智能体导航与多模态推理的学者参考

视觉-语言导航(VLN)旨在使具身智能体根据自然语言指令在真实环境中到达目标位置。以往方法多依赖全局场景表征或物体级特征,难以充分捕捉多模态间的复杂交互。本文提出多层级融合与推理架构(MFRA),通过分层融合机制,整合从低级视觉线索到高级语义概念的多层级特征。进一步设计推理模块,利用融合表示进行导航决策,结合指令引导注意力和动态上下文融合。通过选择性捕获并组合相关视觉、语言与时间信号,提升复杂场景下的决策准确性。在REVERIE、R2R和SOON等基准数据集上的实验表明,MFRA性能优于当前最优方法,验证了多层级模态融合在具身导航中的有效性。

原文摘要 · Abstract (English)

Vision-and-Language Navigation (VLN) aims to enable embodied agents to follow natural language instructions and reach target locations in real-world environments. While prior methods often rely on either global scene representations or object-level features, these approaches are insufficient for capturing the complex interactions across modalities required for accurate navigation. In this paper, we propose a Multi-level Fusion and Reasoning Architecture (MFRA) to enhance the agent's ability to reason over visual observations, language instructions and navigation history. Specifically, MFRA introduces a hierarchical fusion mechanism that aggregates multi-level features-ranging from low-level visual cues to high-level semantic concepts-across multiple modalities. We further design a reasoning module that leverages fused representations to infer navigation actions through instruction-guided attention and dynamic context integration. By selectively capturing and combining relevant visual, linguistic, and temporal signals, MFRA improves decision-making accuracy in complex navigation scenarios. Extensive experiments on benchmark VLN datasets including REVERIE, R2R, and SOON demonstrate that MFRA achieves superior performance compared to state-of-the-art methods, validating the effectiveness of multi-level modal fusion for embodied navigation.

视觉导航多模态融合具身智能语言理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。