不重训练也能跨任务复用场景记忆,提升导航成功率。
VTM-Nav: Harnessing Cross-Episode Experience for Object-Goal Navigation with Hierarchical Visual-Topological Memory

- 构建分层视觉-拓扑记忆,按房间结构索引视觉经验。
- 在多个数据集上比无记忆基线高出4.6~5.5的准确率点。
- 适合研究零样本跨会话智能体与长期记忆导航系统。
训练无关的ObjectNav智能体越来越多地使用视觉语言模型(VLMs),但通常在每次请求后丢弃场景知识。本文研究跨会话的ObjectNav,即每个任务独立初始化、单目标,仅保留自建的场景级记忆。我们提出 extit{VTM-Nav},一种无需训练的框架,采用持久的分层视觉-拓扑记忆(VTM)。VTM利用粗粒度房间拓扑索引房间内视觉记忆,区分室内与远距离可见证据,并保留成功接近线索。每项任务中,VTM-Nav基于累积场景结构重新定位智能体,从可能房间检索目标相关记录,并将记忆引导锚定于当前观测推导的候选区域。保守执行保护机制应对局部失败。在相同40步条件下,VTM-Nav在HM3D v0.1、v0.2和MP3D上分别优于无记忆控制组WMNav 4.6、2.0和0.8的命中率(SR)点,且SPL相当或更高;在HM3D上,也超过文本记忆驱动的WMNav达3.1和5.5的SR点。结果证明通过分层视觉-拓扑记忆可有效复用跨会话场景经验。
原文摘要 · Abstract (English)
Training-free ObjectNav agents increasingly use vision-language models (VLMs), yet typically discard acquired scene knowledge after each request. We study cross-episode ObjectNav, where each request is an independently initialized, single-goal episode and only self-acquired, scene-scoped memory persists across episodes. We ask whether an agent with fixed model parameters and navigation components can reuse such experience without retraining or oracle information. We introduce \method, a training-free framework with a persistent Hierarchical Visual-Topological Memory (VTM). VTM uses a coarse room topology to index room-owned visual memories, distinguishes in-room from remote-visible evidence, and retains successful approach cues. For each request, VTM-Nav re-localizes the agent in accumulated scene structure, retrieves target-relevant records from plausible rooms, and grounds memory guidance in candidates derived from the current observation. A conservative execution guard further handles local failures. Under matched 40-step comparisons, VTM-Nav exceeds the memory-reset WMNav control by 4.6, 2.0, and 0.8 SR points on HM3D v0.1, HM3D v0.2, and MP3D, respectively, with comparable or higher SPL. On HM3D, it also exceeds WMNav harnessed by textual memory by 3.1 and 5.5 SR points. These results demonstrate effective reuse of cross-episode scene experience through hierarchical visual-topological memory.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。