arXiv:2503.07323cs.AIcs.CV2025-03被引 3

用大模型实现多智能体动态环境下的自主导航

Navigating Motion Agents in Dynamic and Cluttered Environments through LLM Reasoning

  • 将环境、智能体和路径统一编码为语言符号,让大模型做空间推理
  • 无需训练即可支持多智能体协作与动态避障,可跨场景通用
  • 适合研究交互式机器人导航或具身智能的学者与开发者

本文推动基于大语言模型(LLMs)的运动智能体在动态复杂环境中的自主导航,显著超越以往仅限于静态简单环境、单智能体且仅四方向移动的开创性研究。我们探索将LLM作为空间推理器,将真实室内布局、可动态移动的智能体及其路径统一编码为离散符号,类似语言标记。所提出的无训练框架支持多智能体协同、闭环重规划及动态障碍物避让,无需重新训练或微调。实验表明,仅通过文本交互,LLM可在不同智能体、任务与环境中实现泛化,为模拟与具身系统中语义驱动的交互式导航开辟新可能。

原文摘要 · Abstract (English)

This paper advances motion agents empowered by large language models (LLMs) toward autonomous navigation in dynamic and cluttered environments, significantly surpassing first and recent seminal but limited studies on LLM's spatial reasoning, where movements are restricted in four directions in simple, static environments in the presence of only single agents much less multiple agents. Specifically, we investigate LLMs as spatial reasoners to overcome these limitations by uniformly encoding environments (e.g., real indoor floorplans), agents which can be dynamic obstacles and their paths as discrete tokens akin to language tokens. Our training-free framework supports multi-agent coordination, closed-loop replanning, and dynamic obstacle avoidance without retraining or fine-tuning. We show that LLMs can generalize across agents, tasks, and environments using only text-based interactions, opening new possibilities for semantically grounded, interactive navigation in both simulation and embodied systems.

智能体导航大模型动态避障多智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。