多专家模型协作让机器人任务规划更可靠。
HEART: Coordination of Heterogeneous Expert Agents for Physically Grounded Robotic Task Planning

- 分任务给不同专长的模型,按角色分工推理。
- 在真实场景下成功率提升,比单模型高30%以上。
- 适合资源有限的机器人系统,部署更高效。
大型语言模型(LLMs)能处理复杂指令,但在机器人任务规划中常忽略物理和空间约束。现有基于LLM的规划器直接将文本转为动作序列,缺乏对可行性、可达性和逻辑顺序的结构化推理,导致计划无效或不完整。我们提出一种异构多LLM框架,将指令分解为原子推理任务,并在令牌预算下分配给角色特化的专家代理,以应对现实中的计算与通信约束。通过结合角色导向推理与约束驱动的计划合成,HEART在规划前验证能力、可达性及约束条件,生成可物理执行的计划,同时保持效率。在多个家庭基准测试中,HEART持续优于单LLM与基于规则的规划器,证明异构LLM协作可在资源受限下实现鲁棒且可扩展的机器人任务规划。
原文摘要 · Abstract (English)
Large Language Models (LLMs) can reason over complex instructions but often fail to satisfy the physical and spatial constraints required for robotic task planning. Recent LLM-based planners directly translate text into action sequences, yet they lack structured reasoning about feasibility, reachability, and logical order, resulting in invalid or incomplete plans. We present a heterogeneous multi-LLM framework that decomposes instructions into atomic reasoning tasks and allocates them to role-specialized expert agents under a token budget for real-world computational and communicational constraints. By combining role-oriented reasoning from heterogeneous agents followed by constraint-driven plan synthesis, HEART validates capability, reachability, and constraint conditions before planning and helps produce physically executable plans while maintaining efficiency. Experiments across different household benchmarks show that HEART consistently improves plan success compared to single-LLM and rule-based planners, demonstrating that heterogeneous LLM collaboration enables robust and scalable robotic task planning under resource constraints.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。