arXiv:2607.13072cs.ROcs.AI2026-07中稿 · publication by IEE…

用分层空间认知让AI零样本找物更准,像人一样先看房间再找东西。

HRO: Hierarchical Room-to-Object Framework for Zero-Shot Object Goal Navigation with Large Language Models

论文配图:HRO: Hierarchical Room-to-Object Framework for Zero-Shot Object Goal Navigation with Large Language Models
图 1 · 摘自论文原文
  • 构建房间到物体的分层导航框架,从宏观到微观逐步定位目标。
  • 在Gibson和HM3D数据集上成功率达90%以上,显著优于现有方法。
  • 适合研究智能体导航、大模型推理与场景理解的学者参考。

零样本物体目标导航旨在使智能体在未见过的环境中探索并导航至未知类别的物体,无需针对特定目标进行训练。现有基于大语言模型(LLMs)的方法通常将LLM作为扁平化推理工具,直接关联物体或区域,缺乏类人级的房间语义层次化空间认知,导致探索盲区大、语义关联精度低,且未能充分发挥LLM的常识推理潜力。本文提出一种由大语言模型驱动的分层房间到物体(HRO)框架,引导智能体以粗粒度到细粒度的方式探索并导航至目标物体。在Gibson和HM3D数据集上的实验表明,所提HRO框架在成功率和泛化能力方面均优于现有基于LLM的方法,凸显了大语言模型在零样本物体目标导航中的强大潜力。

原文摘要 · Abstract (English)

Zero-shot object-goal navigation aims to enable an intelligent agent to explore and navigate to objects of unknown categories in an unfamiliar environment without specific target training. In zero-shot navigation tasks, pre-trained large models are usually employed to leverage their prior knowledge for guiding the agent's navigation. However, existing zero-shot object-goal navigation methods based on large language models (LLMs) merely utilize LLMs as flat reasoning tools to directly associate objects or regions. They lack the hierarchical spatial cognition modeling of human-like room semantics to object localization, which leads to strong blindness in exploration, insufficient accuracy in semantic association, and failure to fully unleash the common-sense reasoning potential of LLMs. This paper proposes an LLM-driven hierarchical room-to-object (HRO) framework for zero-shot object-goal navigation, which guides the agent to explore and navigate to the target object in a coarse-to-fine manner. Experiments on Gibson and HM3D datasets verify that our HRO framework achieves superior success rate and generalization over existing LLM-based methods, underscoring LLMs' strong potential for zero-shot object-goal navigation.

零样本导航大模型分层推理智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。