arXiv:2608.13923cs.CVcs.RO2026-08

提出可保留多假设的物体记忆,提升导航中语义理解的鲁棒性。

OpenBelief-Nav: Evidence-Preserving Object Memory for Open-Vocabulary Language-Guided Navigation

论文配图:OpenBelief-Nav: Evidence-Preserving Object Memory for Open-Vocabulary Language-Guided Navigation
图 1 · 摘自论文原文
  • 保留观测级短语与置信度,不提前锁定语义标签
  • 在ScanNet200上实现0.2742的mIoU,优于基线0.2393
  • 适合需要多视角推理的开放词汇导航任务

开放词汇3D场景图为语言引导导航提供了紧凑的语义记忆,但物体常被融合为单一特征或固定标签,导致少数但任务相关的假设丢失。我们提出OpenBelief-Nav,一种保留观测级短语、可靠性线索和帧掩码来源的证据保全型物体记忆,同时维护独立的几何与视觉表征。语义相关短语被整合为与词汇无关的物体信念,支持固定词汇投影或自由形式检索。在五个ScanNet200和八个Replica场景中,完整信念投影的mIoU分别达0.2742和0.2912,高于匹配的早期承诺读取方法(0.2393和0.2701)。在78次HM3D-YCB导航试验中,共识与早期承诺检索各成功60/78次,而信念加权检索为58/78,DualMap为55/78。在20次Unitree G1运行(10组匹配评估)中,允许最多两次验证候选的修正策略将目标确认成功率从6/10提升至8/10。代码将在录用后公开于https://openbelief-nav.github.io/。

原文摘要 · Abstract (English)

Open-vocabulary 3D scene graphs provide compact semantic memory for language-guided navigation, but mapped objects are often exposed through a single fused feature or committed semantic label. Such commitment can remove minority yet task-relevant hypotheses from the task-time interface. We present OpenBelief-Nav, an evidence-preserving object memory that retains observation-level phrases, reliability cues, and frame-mask provenance while maintaining separate aggregate geometric and visual representations. Semantically related phrases are consolidated into a vocabulary-independent object belief from which task-specific readouts perform fixed-vocabulary projection or free-form retrieval. On five ScanNet200 and eight Replica scenes, full-belief projection achieves mIoU scores of 0.2742 and 0.2912, compared with 0.2393 and 0.2701 for a matched early-commit readout. Across 78 HM3D-YCB navigation trials, consensus and early-commit retrieval each achieve 60/78 successes, compared with 58/78 for belief-weighted retrieval and 55/78 for DualMap. Across 20 Unitree G1 runs organized as 10 matched evaluation cases, a correction policy permitting at most two verified candidate attempts improves target-confirmation success from 6/10 to 8/10 relative to top-1-only execution. Code will be released upon acceptance at https://openbelief-nav.github.io/.

导航语义记忆开放词汇多假设

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。