arXiv:2607.25448cs.RO2026-07中稿 · IROS 2026

用房间语义间接推断物体共现,实现无需训练的零样本导航

Room-Mediated Co-occurrence for Zero-Shot Object-Centric Semantic Navigation via Frontier Scoring

论文配图:Room-Mediated Co-occurrence for Zero-Shot Object-Centric Semantic Navigation via Frontier Scoring
图 1 · 摘自论文原文
  • 通过房间词典将物体映射为概率向量,间接计算共现关系
  • 在HM3D上提升成功率3%、加权路径长度1.3%,且无需训练
  • 适合需要可解释性与开放词汇的零样本导航研究者

零样本物体导航方法越来越多地依赖视觉-语言先验,但潜在空间中的物体-物体相似性往往无法有效反映空间共现。本文提出一种分析性、无需训练的语义导航流程,通过紧凑的房间词典中介物体关系。每个物体标签被映射为基于CLIP的房间概率向量(RPV),物体目标共现度由RPV分布重叠计算。这些得分通过测地线洪水填充传播(快速行进法),结合自适应信号衰减,投影到价值图中,并用于根据语义得分对前沿进行排序以指导导航。上述组件构成一个集成的、无需训练的、以物体为中心的开放词汇零样本导航管道。实验表明,在HM3D数据集验证集上,相比图像整体基线,本方法相对提升了3%的成功率(SR)和1.3%的加权逆路径长度成功率(SPL),同时保持可解释性和开放词汇灵活性。代码已公开:uts-ri.github.io/RPV-SemNav。

原文摘要 · Abstract (English)

Zero-shot ObjectNav methods increasingly use vision-language priors, but direct object-object similarity in the latent space is often a weak proxy for spatial co-occurrence. We present an analytical, training-free semantic navigation pipeline that mediates object relationships through a compact room lexicon. Each object label is mapped to a CLIP-derived Room Probability Vector (RPV), and object-target co-occurrence is computed from RPV distribution overlap. These scores are projected onto a value map using geodesic flood-fill propagation (Fast Marching Method), with adaptive signal decay, and are used to rank frontiers by semantic score for navigation. Together, these components form an integrated, training-free, object-centric pipeline for open-vocabulary zero-shot navigation. Results show that our object-centric approach improves Success Rate (SR) and Success by weighted inverse Path Length (SPL) by a relative 3% and 1.3%, respectively, compared to image-holistic baselines on the HM3D dataset validation split, while preserving interpretability and open-vocabulary flexibility. Code is available at: uts-ri.github.io/RPV-SemNav.

零样本导航物体中心语义推理开放词汇

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。