系统梳理大模型规划能力,为智能体工作流提供评估框架。
PlanGenLLMs: A Modern Survey of LLM Planning Capabilities
- 从六项核心标准分析大模型规划方法
- 对比不同任务中规划系统的优劣表现
- 适合研究智能体与自动化决策的开发者参考
大语言模型在生成计划方面具有巨大潜力,可将初始状态转化为目标状态。大量研究探索了其在网页导航、旅行规划和数据库查询等任务中的应用,但多数系统针对特定问题设计,难以比较或迁移。同时,缺乏统一的评估标准。本综述基于经典规划理论,系统考察完整性、可执行性、最优性、表示能力、泛化性和效率六项关键指标,分析代表性成果的优缺点,并指出未来发展方向,为从业者和初学者提供实用指导。
原文摘要 · Abstract (English)
LLMs have immense potential for generating plans, transforming an initial world state into a desired goal state. A large body of research has explored the use of LLMs for various planning tasks, from web navigation to travel planning and database querying. However, many of these systems are tailored to specific problems, making it challenging to compare them or determine the best approach for new tasks. There is also a lack of clear and consistent evaluation criteria. Our survey aims to offer a comprehensive overview of current LLM planners to fill this gap. It builds on foundational work by Kartam and Wilkins (1990) and examines six key performance criteria: completeness, executability, optimality, representation, generalization, and efficiency. For each, we provide a thorough analysis of representative works and highlight their strengths and weaknesses. Our paper also identifies crucial future directions, making it a valuable resource for both practitioners and newcomers interested in leveraging LLM planning to support agentic workflows.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。