arXiv:2604.00265cs.CVcs.AI2026-04被引 1

首个可复现的协作导航基准,分离评估对话与导航能力。

Benchmarking Interaction, Beyond Policy: a Reproducible Benchmark for Collaborative Instance Object Navigation

论文配图:Benchmarking Interaction, Beyond Policy: a Reproducible Benchmark for Collaborative Instance Object Navigation
图 1 · 摘自论文原文
  • 设计独立评分的提问协议,分离评估交互与导航。
  • 提供2.8万条高质量对话轨迹,支持模型训练与分析。
  • 轻量模型比现有方法小3倍快70倍,泛化能力更强。

我们提出Question-Asking Navigation (QAsk-Nav),首个可复现的协作实例对象导航(CoIN)基准,能够对具身导航与协作提问进行显式、独立评估。CoIN任务要求智能体在部分可观测环境下,仅通过本体视觉与人类进行自然语言对话,定位自由格式描述的目标对象,对话可帮助区分视觉相似的物体实例。现有CoIN基准主要关注导航成功率,缺乏对协作交互的一致评估支持。QAsk-Nav通过三项改进:(i) 轻量级提问协议,独立于导航评分;(ii) 增强导航协议,包含真实、多样、高质量的目标描述;(iii) 开源数据集,包含28,000条经质量验证的推理与提问轨迹,用于训练和分析CoIN模型的交互能力。基于该基准,我们构建Light-CoNav,一种轻量级统一模型,相较现有模块化方法体积缩小3倍、速度提升70倍,且在未见物体与环境上的泛化性能优于当前最优方法。

原文摘要 · Abstract (English)

We propose Question-Asking Navigation (QAsk-Nav), the first reproducible benchmark for Collaborative Instance Object Navigation (CoIN) that enables an explicit, separate assessment of embodied navigation and collaborative question asking. CoIN tasks an embodied agent with reaching a target specified in free-form natural language under partial observability, using only egocentric visual observations and interactive natural-language dialogue with a human, where the dialogue can help to resolve ambiguity among visually similar object instances. Existing CoIN benchmarks are primarily focused on navigation success and offer no support for consistent evaluation of collaborative interaction. To address this limitation, QAsk-Nav provides (i) a lightweight question-asking protocol scored independently of navigation, (ii) an enhanced navigation protocol with realistic, diverse, high-quality target descriptions, and (iii) an open-source dataset, that includes 28,000 quality-checked reasoning and question-asking traces for training and analysis of interactive capabilities of CoIN models. Using the proposed QAsk-Nav benchmark, we develop Light-CoNav, a lightweight unified model for collaborative navigation that is 3x smaller and 70x faster than existing modular methods, while outperforming state-of-the-art CoIN approaches in generalization to unseen objects and environments. Project page at https://benchmarking-interaction.github.io/

协作导航对话评估基准测试轻量模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。