arXiv:2503.02247cs.CVcs.RO2025-03被引 63

用视觉语言模型构建可预测的环境世界模型,提升导航成功率与探索效率。

WMNav: Integrating Vision-Language Models into World Models for Object Goal Navigation

  • 基于视觉语言模型构建可动态更新的世界模型,预测决策后果并生成记忆反馈。
  • 在HM3D和MP3D上实现+3.2%成功率和+13.5%成功率提升,优于现有零样本基准。
  • 适合关注具身智能、目标导航与多模态规划的研究者参考。

物体目标导航——要求智能体在未见过的环境中定位特定物体——仍是具身人工智能的核心挑战。尽管基于视觉语言模型(VLM)的智能体在提示驱动下展现出良好的感知与决策能力,但尚无系统建立完全模块化的世界模型设计,以通过预测世界未来状态来减少与环境的高风险、高成本交互。本文提出WMNav,一种基于视觉语言模型(VLM)的新型世界模型导航框架。它能预测决策可能结果,并构建记忆以向策略模块提供反馈。为保留环境预测状态,WMNav引入在线维护的惊奇值地图作为世界模型记忆的一部分,为导航策略提供动态配置。通过模仿人类思维过程分解决策,WMNav利用世界模型计划与观测间的反馈差异有效缓解模型幻觉影响。为进一步提升效率,采用两阶段动作提议策略:广域探索后接精确定位。在HM3D和MP3D上的大量评估表明,WMNav在成功率(+3.2% SR,+13.5% SR)与探索效率(+3.2% SPL,+1.1% SPL)上均超越现有零样本基准。

原文摘要 · Abstract (English)

Object Goal Navigation-requiring an agent to locate a specific object in an unseen environment-remains a core challenge in embodied AI. Although recent progress in Vision-Language Model (VLM)-based agents has demonstrated promising perception and decision-making abilities through prompting, none has yet established a fully modular world model design that reduces risky and costly interactions with the environment by predicting the future state of the world. We introduce WMNav, a novel World Model-based Navigation framework powered by Vision-Language Models (VLMs). It predicts possible outcomes of decisions and builds memories to provide feedback to the policy module. To retain the predicted state of the environment, WMNav proposes the online maintained Curiosity Value Map as part of the world model memory to provide dynamic configuration for navigation policy. By decomposing according to a human-like thinking process, WMNav effectively alleviates the impact of model hallucination by making decisions based on the feedback difference between the world model plan and observation. To further boost efficiency, we implement a two-stage action proposer strategy: broad exploration followed by precise localization. Extensive evaluation on HM3D and MP3D validates WMNav surpasses existing zero-shot benchmarks in both success rate and exploration efficiency (absolute improvement: +3.2% SR and +3.2% SPL on HM3D, +13.5% SR and +1.1% SPL on MP3D). Project page: https://b0b8k1ng.github.io/WMNav/.

目标导航世界模型视觉语言模型具身智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。