arXiv:2603.26788cs.ROcs.CV2026-03被引 2

用记忆增强框架解决未知物体导航难题,提升探索效率与准确性。

ReMemNav: A Rethinking and Memory-Augmented Framework for Zero-Shot Object Navigation

  • 引入全景语义先验与情景记忆,融合视觉语言模型进行分层导航。
  • 在HM3D和MP3D上成功率提升最高达18.2%,探索效率显著改善。
  • 适合研究零样本导航、具身智能与记忆增强决策的学者参考。

零样本物体导航要求智能体在无先验地图或任务特定训练的情况下,定位未见目标物体,仍面临重大挑战。尽管视觉语言模型(VLMs)具备潜在常识推理能力,但仍存在空间幻觉、局部探索死锁以及高层语义意图与底层控制脱节等问题。为此,我们提出新型分层导航框架ReMemNav,无缝融合全景语义先验与情景记忆。引入Recognize Anything Model锚定VLM的空间推理过程,并设计基于情景语义缓冲队列的自适应双模态重思机制,主动验证目标可见性并利用历史记忆修正决策以避免死锁。低层动作执行中,通过深度掩码提取可行动作序列,使VLM可选择最优动作映射为实际空间移动。在HM3D与MP3D上的大量评估表明,ReMemNav优于现有无需训练的零样本基线,在成功率(SR)与路径长度归一化成功度(SPL)上均有显著提升:在HM3D v0.1上分别提升1.7%与7.0%,在HM3D v0.2上提升18.2%与11.1%,在MP3D上提升8.7%与7.9%。

原文摘要 · Abstract (English)

Zero-shot object navigation requires agents to locate unseen target objects in unfamiliar environments without prior maps or task-specific training which remains a significant challenge. Although recent advancements in vision-language models(VLMs) provide promising commonsense reasoning capabilities for this task, these models still suffer from spatial hallucinations, local exploration deadlocks, and a disconnect between high-level semantic intent and low-level control. In this regard, we propose a novel hierarchical navigation framework named ReMemNav, which seamlessly integrates panoramic semantic priors and episodic memory with VLMs. We introduce the Recognize Anything Model to anchor the spatial reasoning process of the VLM. We also design an adaptive dual-modal rethinking mechanism based on an episodic semantic buffer queue. The proposed mechanism actively verifies target visibility and corrects decisions using historical memory to prevent deadlocks. For low-level action execution, ReMemNav extracts a sequence of feasible actions using depth masks, allowing the VLM to select the optimal action for mapping into actual spatial movement. Extensive evaluations on HM3D and MP3D demonstrate that ReMemNav outperforms existing training-free zero-shot baselines in both success rate and exploration efficiency. Specifically, we achieve significant absolute performance improvements, with SR and SPL increasing by 1.7% and 7.0% on HM3D v0.1, 18.2% and 11.1% on HM3D v0.2, and 8.7% and 7.9% on MP3D.

零样本导航记忆增强视觉语言模型具身智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。