用符号规划约束大模型,让机器人任务规划更可靠、可重复。
Constrained Natural Language Action Planning for Resilient Embodied Systems
- 大模型规划+符号系统监督,兼顾灵活性与可靠性
- 仿真环境成功率99%,真实机器人100%成功
- 适合需要高可靠性与透明性的现实机器人应用
在真实世界环境中实现类人智能的具身任务执行仍具挑战性,因环境高度开放且不可控。大型语言模型(LLMs)用于任务规划虽能应对复杂状态/动作空间,但幻觉问题限制其可靠性,且提示工程缺乏透明性与可复现性。而符号规划方法虽可靠可复现,却难以应对现实任务的复杂性与模糊性。本文提出一种新方法:以符号规划对大模型规划进行监督,增强可靠性与可复现性,并通过明确硬约束提供比传统提示工程更强的清晰度。该方法保留大模型的推理能力与开放环境泛化能力。在ALFWorld基准上,本方法达到近完美的99%成功率;部署至真实四足机器人时,任务成功率100%,显著优于纯大模型(50%)与纯符号规划(30%)。该工作为提升基于大模型的机器人规划系统的可靠性、可复现性与透明性提供了有效路径,同时保持其关键优势:灵活与泛化能力。
原文摘要 · Abstract (English)
Replicating human-level intelligence in the execution of embodied tasks remains challenging due to the unconstrained nature of real-world environments. Novel use of large language models (LLMs) for task planning seeks to address the previously intractable state/action space of complex planning tasks, but hallucinations limit their reliability, and thus, viability beyond a research context. Additionally, the prompt engineering required to achieve adequate system performance lacks transparency, and thus, repeatability. In contrast to LLM planning, symbolic planning methods offer strong reliability and repeatability guarantees, but struggle to scale to the complexity and ambiguity of real-world tasks. We introduce a new robotic planning method that augments LLM planners with symbolic planning oversight to improve reliability and repeatability, and provide a transparent approach to defining hard constraints with considerably stronger clarity than traditional prompt engineering. Importantly, these augmentations preserve the reasoning capabilities of LLMs and retain impressive generalization in open-world environments. We demonstrate our approach in simulated and real-world environments. On the ALFWorld planning benchmark, our approach outperforms current state-of-the-art methods, achieving a near-perfect 99% success rate. Deployment of our method to a real-world quadruped robot resulted in 100% task success compared to 50% and 30% for pure LLM and symbolic planners across embodied pick and place tasks. Our approach presents an effective strategy to enhance the reliability, repeatability and transparency of LLM-based robot planners while retaining their key strengths: flexibility and generalizability to complex real-world environments. We hope that this work will contribute to the broad goal of building resilient embodied intelligent systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。