arXiv:2607.15758cs.RO2026-07

通过评分层技能干预,让视觉语言导航模型零样本提升路径规划能力。

SkillNav: Score-Level Skill Intervention for Zero-Shot Object Goal Navigation

论文配图:SkillNav: Score-Level Skill Intervention for Zero-Shot Object Goal Navigation
图 1 · 摘自论文原文
  • 在现有模型评分图上叠加可组合的三层行为技能,零成本记录导航记忆。
  • 在多个数据集上达到新SOTA,HM3D v0.2 SPL达43.2,较前人提升6.0点。
  • 无需训练,适合希望快速增强导航模型行为策略的研究者与开发者。

视觉语言模型(VLM)代理在零样本物体目标导航任务中取得进展,但单帧推理使其缺乏跨步的行为感知,导致反复出现死胡同停滞、室内循环和绕远路径等失败。基于提示的修复方法会增加多模块任务中的令牌开销,仍难以编码角度、地图单元格和视角坐标等空间信号。本文提出SkillNav,一种可扩展的VLM导航行为技能框架,将现代VLM导航器已维护的趣味性价值图作为可写载体,在不消耗额外令牌的情况下,通过可组合技能在其中注入行为记忆。技能按行为权限分为三层:软缩放(比例重加权)、下界提升(区域级保障)和硬覆盖(阈值触发强制动作),并按固定顺序协作,确立明确优先级。该设计使能力提升转化为技能注册:新行为可插拔式接入,无需重训练VLM或干扰已有技能,实现持续优化。一个最小提示通道补充类别级语义提示,形成双表示架构:空间记忆保留在地图中,语义记忆存储于短提示。训练免费,SkillNav在MP3D(SPL 25.5)、HM3D v0.1(SPL 39.3)、HM3D v0.2(SPL 43.2)上建立新SOTA,相比最强基线提升最高达6.0绝对点,并在HM3D v0.1(成功率69.7)和v0.2(成功率75.9)上取得最高成功率达。

原文摘要 · Abstract (English)

Vision-Language Model (VLM) agents have advanced zero-shot object-goal navigation, yet single-frame reasoning leaves them without the cross-step behavioral awareness an embodied navigator requires, producing recurring failures such as dead-end stalls, in-room loops, and circuitous approaches to detected targets. Prompt-based remedies inflate token budgets across multi-submodule episodes and still struggle to encode inherently spatial signals such as angles, map cells, and viewpoint coordinates. In this paper, we propose SkillNav, an extensible behavioral skill framework for VLM-based navigation that treats the curiosity value map already maintained by modern VLM navigators as a writable substrate on which composable skills inscribe behavioral memory at zero token cost. Skills are stratified into three tiers by their level of behavioral authority, namely soft scaling for proportional reweighting, lower-bound boost for region-level guarantees, and hard override for threshold-triggered forced actions, and cooperate across tiers under a fixed composition order that establishes a predictable, declared priority among skills. This design turns capability improvement into skill registration: new behaviors plug in without retraining the VLM or disturbing existing skills, opening a path for continual refinement. A minimal prompt channel complements the score-level skills with category-level semantic hints, yielding a dual-representation design in which spatial memory lives on the map and semantic memory in short prompts. Training-free, SkillNav establishes new state-of-the-art SPL across MP3D (25.5), HM3D v0.1 (39.3), and HM3D v0.2 (43.2), improving SPL by up to 6.0 absolute over the strongest prior method, and achieves the highest Success Rate on HM3D v0.1 (69.7) and v0.2 (75.9).

导航视觉语言模型技能框架零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。