用PDDL增强的LLM集群,让多模态任务自动规划执行
From An LLM Swarm To A PDDL-Empowered HIVE: Planning Self-Executed Instructions In A Multi-Modal Jungle
- 基于LLM与PDDL构建可解释任务规划框架
- 支持多模态输入输出,实现复杂任务链自动调度
- 通过MuSE基准测试验证性能与透明性优势
针对深度模型生态日益增强能力带来的代理系统需求,我们提出Hive——一种知识感知的任务规划系统,能根据自然语言指令调度并执行原子动作序列,自动选择合适模型完成任务。Hive在多个模型集合上运行,支持多模态输入输出,可在满足用户约束条件下,实现复杂真实场景查询的可解释规划。系统采用基于LLM的正式逻辑架构,结合PDDL操作保障规划过程透明性。为全面评估多模态代理系统能力,我们引入MuSE基准。实验表明,该框架在跨模型任务选择上达到新基准,显著优于现有方法,同时确保可解释性并严格遵循用户约束。
原文摘要 · Abstract (English)
In response to the call for agent-based solutions that leverage the ever-increasing capabilities of the deep models' ecosystem, we introduce Hive -- a comprehensive solution for knowledge-aware planning of a set of atomic actions to address input queries and subsequently selecting appropriate models accordingly. Hive operates over sets of models and, upon receiving natural language instructions (i.e. user queries), schedules and executes explainable plans of atomic actions. These actions can involve one or more of the available models to achieve the overall task, while respecting end-users specific constraints. Notably, Hive handles tasks that involve multi-modal inputs and outputs, enabling it to handle complex, real-world queries. Our system is capable of planning complex chains of actions while guaranteeing explainability, using an LLM-based formal logic backbone empowered by PDDL operations. We introduce the MuSE benchmark in order to offer a comprehensive evaluation of the multi-modal capabilities of agent systems. Our findings show that our framework redefines the state-of-the-art for task selection, outperforming other competing systems that plan operations across multiple models while offering transparency guarantees while fully adhering to user constraints.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。