arXiv:2505.20828cs.RO2025-05被引 2

让机器人在未知大环境里高效找物,靠的是智能推理与经验积累结合。

GET: Goal-directed Exploration and Targeting for Large-Scale Unknown Environments

  • 用角色反馈机制实现实时决策,融合任务目标与记忆信息。
  • 通过高斯混合模型构建概率任务地图,持续更新物体位置先验。
  • 在真实大规模环境中显著优于传统方法,适合复杂动态场景应用。

在大型非结构化环境中进行物体搜索仍是机器人领域的基本挑战,尤其在动态或广阔场景如户外自主探索中。该任务需要强大的空间推理能力以及利用过往经验的能力。尽管大语言模型(LLMs)具备出色的语义理解能力,但其在具身任务中的应用受限于空间推理的“接地”缺陷及记忆整合与决策一致性不足的问题。为此,我们提出GET(目标导向探索与定位)框架,通过将基于LLM的推理与经验引导探索相结合,提升物体搜索性能。核心是DoUT(统一思维图)推理模块,通过角色化反馈环实现实时决策,整合任务特定标准与外部记忆。对于重复任务,GET基于高斯混合模型维护概率任务地图,随环境变化持续更新物体位置先验。在真实世界大规模环境中的实验表明,GET在多种LLM和任务设置下均显著提升搜索效率与鲁棒性,明显优于启发式方法与仅使用LLM的基线。结果表明,结构化地集成LLM可为复杂环境中的具身决策提供可扩展且通用的解决方案。

原文摘要 · Abstract (English)

Object search in large-scale, unstructured environments remains a fundamental challenge in robotics, particularly in dynamic or expansive settings such as outdoor autonomous exploration. This task requires robust spatial reasoning and the ability to leverage prior experiences. While Large Language Models (LLMs) offer strong semantic capabilities, their application in embodied contexts is limited by a grounding gap in spatial reasoning and insufficient mechanisms for memory integration and decision consistency.To address these challenges, we propose GET (Goal-directed Exploration and Targeting), a framework that enhances object search by combining LLM-based reasoning with experience-guided exploration. At its core is DoUT (Diagram of Unified Thought), a reasoning module that facilitates real-time decision-making through a role-based feedback loop, integrating task-specific criteria and external memory. For repeated tasks, GET maintains a probabilistic task map based on a Gaussian Mixture Model, allowing for continual updates to object-location priors as environments evolve.Experiments conducted in real-world, large-scale environments demonstrate that GET improves search efficiency and robustness across multiple LLMs and task settings, significantly outperforming heuristic and LLM-only baselines. These results suggest that structured LLM integration provides a scalable and generalizable approach to embodied decision-making in complex environments.

机器人大模型探索空间推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。