用分层任务网络提升大模型智能体的规划能力
Procedural Knowledge Improves Agentic LLM Workflows
- 用分层任务网络(HTN)结构化地组织任务执行流程
- 200亿参数模型经优化后超越1200亿参数基线模型
- 人工编写或大模型生成的流程知识均有效提升性能
大语言模型在缺乏工具支持、提示工程或微调的情况下,往往难以完成智能体任务。尽管已有研究表明领域相关的程序性知识能显著提升规划效率,但其在需要隐式规划的智能体任务中的潜力仍鲜有评估。本文形式化、实现并评估了一种利用分层任务网络(HTN)的智能体大模型工作流。实证结果表明,手工编写的HTN可大幅改善大模型在智能体任务上的表现;使用HTN后,200亿和700亿参数的模型性能超过1200亿参数的基线模型。此外,由大模型生成的HTN也能提升整体性能,但效果较弱。结果表明,通过人类、文档或大模型提炼程序性知识,将成为提升大模型工作流的重要手段。
原文摘要 · Abstract (English)
Large language models (LLMs) often struggle when performing agentic tasks without substantial tool support, prom-pt engineering, or fine tuning. Despite research showing that domain-dependent, procedural knowledge can dramatically increase planning efficiency, little work evaluates its potential for improving LLM performance on agentic tasks that may require implicit planning. We formalize, implement, and evaluate an agentic LLM workflow that leverages procedural knowledge in the form of a hierarchical task network (HTN). Empirical results of our implementation show that hand-coded HTNs can dramatically improve LLM performance on agentic tasks, and using HTNs can boost a 20b or 70b parameter LLM to outperform a much larger 120b parameter LLM baseline. Furthermore, LLM-created HTNs improve overall performance, though less so. The results suggest that leveraging expertise--from humans, documents, or LLMs--to curate procedural knowledge will become another important tool for improving LLM workflows.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。