arXiv:2608.10817cs.RO2026-08

提出新方法让机器人在陌生环境快速准确定位任意物体。

AECNav: Active Evidence Consolidation for Efficient Zero-Shot Open-Vocabulary Object Navigation

论文配图:AECNav: Active Evidence Consolidation for Efficient Zero-Shot Open-Vocabulary Object Navigation
图 1 · 摘自论文原文
  • 用证据驱动思路统一感知与决策,减少冗余计算。
  • 在三个数据集上最高达84.7%成功率,物理机器人实测95%成功。
  • 适合追求高效零样本导航的机器人研究者使用。

在开放词汇场景下的零样本目标导航(ZSON)极具挑战性,要求机器人在未见过的环境中定位任意指定物体且无需特定任务训练。当前方法因感知流程冗余和证据不足,仍存在延迟高、准确率低的问题。本文将ZSON重构为以证据为核心的感知-决策问题,提出AECNav——一种无需训练的流水线,包含三部分:i) 证据门控感知,通过共享编码建立统一语义基础,消除重复计算;ii) 证据融合,将检测结果聚合为簇级对数似然信念,明确区分真实目标支持与视觉相似干扰物带来的虚假信心,同时将预期检测缺失视为负证据;iii) 主动证据获取,基于信息增益最大化并最小化移动成本,选择前沿区域进行高效探索。实验表明,AECNav显著优于现有方法,在HM3D-v2、HM3D-OVON和MP3D上分别取得84.7%、57.3%和51.3%的成功率,推理开销大幅降低,并在真实四足机器人上实现40次试验中95%的成功率,运行频率约5Hz。代码将在接受后公开。

原文摘要 · Abstract (English)

Zero-shot object-goal navigation (ZSON) in open-vocabulary scenarios is challenging, as it requires a robot to locate an arbitrarily specified object in an unseen environment without task-specific training. Currently, the task still suffers from high latency and limited accuracy due to redundant perception pipelines and insufficient evidence for reliable target confirmation. In this letter, we reframe ZSON as an evidence-driven perception-to-decision problem and present AECNav, a training-free pipeline built on three components: i) Evidence-gated perception, which utilizes a shared encoding across all reasoning stages to establish a unified semantic basis and eliminate redundant computations; ii) Evidence consolidation, which aggregates detections into cluster-level log-odds beliefs. This explicitly separates genuine target support from the false confidence of visually similar distractors, while treating the absence of expected detections as negative evidence; and iii) Active evidence acquisition, which sustains productive exploration under weak semantic cues by selecting frontiers that maximize information gain at minimal traversal cost. As a result, AECNav significantly outperforms previous methods and achieves state-of-the-art success rates of 84.7%, 57.3%, and 51.3% on HM3D-v2, HM3D-OVON, and MP3D, respectively, with substantially lower inference overhead, and attains 95% success across 40 trials on a physical quadruped robot at roughly 5Hz. Code will be made publicly available upon acceptance.

机器人导航零样本证据融合开放词汇

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。