arXiv:2410.21842cs.CVcs.AI2024-10被引 4

用扩散模型生成未知区域地图,让机器人更聪明地找东西。

Diffusion as Reasoning: Enhancing Object Navigation via Diffusion Model Conditioned on LLM-based Object-Room Knowledge

  • 用扩散模型学习物体在地图中的分布规律,结合已探索区域生成未知区域地图。
  • 在Gibson和MP3D数据集上,导航成功率显著提升,优于现有方法。
  • 利用大模型的常识知识指导物体布局,适合复杂环境下的智能导航研究者。

目标导航(ObjectNav)任务旨在引导智能体在未见过的环境中,通过部分观测找到目标物体。以往方法采用位置预测范式进行长期目标推理,但难以有效整合上下文关系信息;另一类基于地图补全的方法通过生成未探索区域的语义地图来预测长期目标,但未能充分利用已有环境信息,导致地图质量不佳。本文提出一种新方法:训练扩散模型学习语义地图中物体的统计分布模式,并以导航过程中已探索区域的地图为条件,生成未知区域的地图,实现对目标物体的长期目标推理,即“扩散即推理”(DAR)。同时,提出房间引导机制,利用大语言模型(LLM)提取的常识知识,指导扩散模型生成具有房间感知的物体分布。基于生成的地图,智能体将目标物体预测位置设为下一步目标并前进。在Gibson和MP3D数据集上的实验验证了该方法的有效性。

原文摘要 · Abstract (English)

The Object Navigation (ObjectNav) task aims to guide an agent to locate target objects in unseen environments using partial observations. Prior approaches have employed location prediction paradigms to achieve long-term goal reasoning, yet these methods often struggle to effectively integrate contextual relation reasoning. Alternatively, map completion-based paradigms predict long-term goals by generating semantic maps of unexplored areas. However, existing methods in this category fail to fully leverage known environmental information, resulting in suboptimal map quality that requires further improvement. In this work, we propose a novel approach to enhancing the ObjectNav task, by training a diffusion model to learn the statistical distribution patterns of objects in semantic maps, and using the map of the explored regions during navigation as the condition to generate the map of the unknown regions, thereby realizing the long-term goal reasoning of the target object, i.e., diffusion as reasoning (DAR). Meanwhile, we propose the Room Guidance method, which leverages commonsense knowledge derived from large language models (LLMs) to guide the diffusion model in generating room-aware object distributions. Based on the generated map in the unknown region, the agent sets the predicted location of the target as the goal and moves towards it. Experiments on Gibson and MP3D show the effectiveness of our method.

目标导航扩散模型常识推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。