用多智能体架构让大模型机器人更懂协作与应变。
Transforming Monolithic Foundation Models into Embodied Multi-Agent Architectures for Human-Robot Collaboration
- 将大模型拆解为感知、规划、决策等专用智能体协同工作
- 在真实场景中实现90%以上任务成功率,支持动态人机分工
- 适合需要长期稳定服务的机器人团队部署
基础模型虽在机器人感知与规划中发挥核心作用,但其单体设计难以适应实际服务流程的分布式与动态性。视觉语言模型具备强语义理解,却缺乏具身动作能力,依赖人工设计技能;视觉-语言-动作策略虽能实现反应式操作,但在不同机器人形态间表现脆弱,几何定位能力弱,且缺少主动协作机制。这些局限表明,单纯扩大单一模型无法实现人群环境中可靠自主。为此,我们提出InteractGen——一个基于大语言模型的多智能体框架,将机器人智能分解为持续感知、依赖关系感知规划、决策与验证、失败反思及动态人类委托等专用智能体,使基础模型成为闭环集体中的受控组件。该系统在异构机器人团队上部署,并通过为期三个月的开放使用研究评估,显著提升任务成功率、适应性与人机协作水平,证明多智能体协同是实现社会情境下服务自主更可行的路径,而非继续扩展单一模型。
原文摘要 · Abstract (English)
Foundation models have become central to unifying perception and planning in robotics, yet real-world deployment exposes a mismatch between their monolithic assumption that a single model can handle all cognitive functions and the distributed, dynamic nature of practical service workflows. Vision-language models offer strong semantic understanding but lack embodiment-aware action capabilities while relying on hand-crafted skills. Vision-Language-Action policies enable reactive manipulation but remain brittle across embodiments, weak in geometric grounding, and devoid of proactive collaboration mechanisms. These limitations indicate that scaling a single model alone cannot deliver reliable autonomy for service robots operating in human-populated settings. To address this gap, we present InteractGen, an LLM-powered multi-agent framework that decomposes robot intelligence into specialized agents for continuous perception, dependency-aware planning, decision and verification, failure reflection, and dynamic human delegation, treating foundation models as regulated components within a closed-loop collective. Deployed on a heterogeneous robot team and evaluated in a three-month open-use study, InteractGen improves task success, adaptability, and human-robot collaboration, providing evidence that multi-agent orchestration offers a more feasible path toward socially grounded service autonomy than further scaling standalone models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。