arXiv:2410.08189cs.CVcs.RO2024-10NeurIPS被引 196

用3D场景图提升大模型零样本导航能力,效果超越监督方法。

SG-Nav: Online 3D Scene Graph Prompting for LLM-based Zero-shot Object Navigation

论文配图:SG-Nav: Online 3D Scene Graph Prompting for LLM-based Zero-shot Object Navigation
图 1 · 摘自论文原文
  • 用3D场景图结构化环境关系,支持大模型推理
  • 在MP3D等平台上准确率超前代方法10%以上
  • 可自纠正感知错误,适合需要可解释性的应用

本文提出一种新的零样本物体导航框架。现有方法仅用空间封闭物体的文本提示,缺乏充分的场景上下文信息。为更好保留环境信息并发挥大模型推理能力,我们采用3D场景图表示观察到的场景,该图以大模型友好的结构编码物体、组别和房间之间的关系。为此设计了分层思维链提示,使大模型可通过遍历节点与边来根据场景上下文推理目标位置。此外,得益于场景图表示,我们进一步设计了重感知机制,使导航框架具备纠正感知误差的能力。在MP3D、HM3D和RoboTHOR环境中进行大量实验,结果表明SG-Nav在所有基准上零样本准确率均超过先前最优方法10%以上,且决策过程具有可解释性。据我们所知,SG-Nav是首个在挑战性MP3D基准上性能超越监督导航方法的零样本方法。

原文摘要 · Abstract (English)

In this paper, we propose a new framework for zero-shot object navigation. Existing zero-shot object navigation methods prompt LLM with the text of spatially closed objects, which lacks enough scene context for in-depth reasoning. To better preserve the information of environment and fully exploit the reasoning ability of LLM, we propose to represent the observed scene with 3D scene graph. The scene graph encodes the relationships between objects, groups and rooms with a LLM-friendly structure, for which we design a hierarchical chain-of-thought prompt to help LLM reason the goal location according to scene context by traversing the nodes and edges. Moreover, benefit from the scene graph representation, we further design a re-perception mechanism to empower the object navigation framework with the ability to correct perception error. We conduct extensive experiments on MP3D, HM3D and RoboTHOR environments, where SG-Nav surpasses previous state-of-the-art zero-shot methods by more than 10% SR on all benchmarks, while the decision process is explainable. To the best of our knowledge, SG-Nav is the first zero-shot method that achieves even higher performance than supervised object navigation methods on the challenging MP3D benchmark.

零样本导航3D场景图大模型推理可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。