多智能体协作框架让机器人更准地听懂长指令走远路
MA-CoNav: A Master-Slave Multi-Agent Framework with Hierarchical Collaboration and Dual-Level Reflection for Long-Horizon Embodied VLN
- 主从架构分工:各智能体专攻感知、规划、执行、记忆
- 真实室内测试中,导航成功率显著高于主流方法
- 适合研究长时序视觉语言导航的开发者参考
视觉语言导航(VLN)旨在使机器人基于复杂语言指令,在陌生环境中完成长距离导航任务。其成功关键在于构建高效的“语言理解—视觉感知—具身执行”闭环。现有方法在复杂长距离任务中常因单一智能体认知过载,导致感知失真和决策漂移。受分布式认知理论启发,本文提出MA-CoNav多智能体协同导航框架,采用“主-从”层级协作架构,将导航所需的感知、规划、执行与记忆功能解耦并分配给专用智能体。具体而言,主智能体负责全局调度,下属智能体组分工协作:观察智能体生成环境描述,规划智能体执行任务分解与动态验证,执行智能体同步进行建图与动作,记忆智能体管理结构化经验。此外,框架引入“局部-全局”双阶段反思机制,动态优化整个导航流程。实验基于Limo Pro机器人采集的真实室内数据集开展,模型全程未进行场景微调。结果表明,MA-CoNav在多个指标上全面优于现有主流VLN方法。
原文摘要 · Abstract (English)
Vision-Language Navigation (VLN) aims to empower robots with the ability to perform long-horizon navigation in unfamiliar environments based on complex linguistic instructions. Its success critically hinges on establishing an efficient ``language-understanding -- visual-perception -- embodied-execution'' closed loop. Existing methods often suffer from perceptual distortion and decision drift in complex, long-distance tasks due to the cognitive overload of a single agent. Inspired by distributed cognition theory, this paper proposes MA-CoNav, a Multi-Agent Collaborative Navigation framework. This framework adopts a ``Master-Slave'' hierarchical agent collaboration architecture, decoupling and distributing the perception, planning, execution, and memory functions required for navigation tasks to specialized agents. Specifically, the Master Agent is responsible for global orchestration, while the Subordinate Agent group collaborates through a clear division of labor: an Observation Agent generates environment descriptions, a Planning Agent performs task decomposition and dynamic verification, an Execution Agent handles simultaneous mapping and action, and a Memory Agent manages structured experiences. Furthermore, the framework introduces a ``Local-Global'' dual-stage reflection mechanism to dynamically optimize the entire navigation pipeline. Empirical experiments were conducted using a real-world indoor dataset collected by a Limo Pro robot, with no scene-specific fine-tuning performed on the models throughout the process. The results demonstrate that MA-CoNav comprehensively outperforms existing mainstream VLN methods across multiple metrics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。