arXiv:2409.14296cs.AIcs.RO2024-09被引 108

构建开放词汇物体导航数据集,让机器人按自然语言找任意物品。

HM3D-OVON: A Dataset and Benchmark for Open-Vocabulary Object Goal Navigation

论文配图:HM3D-OVON: A Dataset and Benchmark for Open-Vocabulary Object Goal Navigation
图 1 · 摘自论文原文
  • 用15000+真实场景物体实例构建开放词汇导航数据集
  • 新模型在复杂环境中导航成功率更高且抗噪声更强
  • 适合研究具身智能与自然语言交互的学者

我们提出Habitat-Matterport 3D开放词汇物体目标导航数据集(HM3D-OVON),这是一个大规模基准,扩展了以往物体导航(ObjectNav)数据集的范围与语义覆盖。基于HM3DSem数据集,该数据集包含来自379个不同类别的超过15,000个标注的家居物体实例,源自真实世界环境的高保真3D扫描。与早期仅限定6-20个预定义类别的数据集不同,HM3D-OVON支持在测试时通过自由形式语言定义目标物体,实现开集目标导航。这一设定推动模型学习能够以开放式词汇搜索文本指定任意物体的视觉语义导航能力。我们系统评估并对比了多种方法,发现可在该数据集上训练出性能更高、对定位与动作噪声更鲁棒的开放词汇导航代理。我们希望该基准及基线结果能促进发展可理解自然语言指令、在真实空间中寻找家居物品的具身智能体,迈向更灵活的人类级语义视觉导航。代码与视频见:naoki.io/ovon。

原文摘要 · Abstract (English)

We present the Habitat-Matterport 3D Open Vocabulary Object Goal Navigation dataset (HM3D-OVON), a large-scale benchmark that broadens the scope and semantic range of prior Object Goal Navigation (ObjectNav) benchmarks. Leveraging the HM3DSem dataset, HM3D-OVON incorporates over 15k annotated instances of household objects across 379 distinct categories, derived from photo-realistic 3D scans of real-world environments. In contrast to earlier ObjectNav datasets, which limit goal objects to a predefined set of 6-20 categories, HM3D-OVON facilitates the training and evaluation of models with an open-set of goals defined through free-form language at test-time. Through this open-vocabulary formulation, HM3D-OVON encourages progress towards learning visuo-semantic navigation behaviors that are capable of searching for any object specified by text in an open-vocabulary manner. Additionally, we systematically evaluate and compare several different types of approaches on HM3D-OVON. We find that HM3D-OVON can be used to train an open-vocabulary ObjectNav agent that achieves both higher performance and is more robust to localization and actuation noise than the state-of-the-art ObjectNav approach. We hope that our benchmark and baseline results will drive interest in developing embodied agents that can navigate real-world spaces to find household objects specified through free-form language, taking a step towards more flexible and human-like semantic visual navigation. Code and videos available at: naoki.io/ovon.

物体导航开放词汇具身智能自然语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。