让语言代理分步思考,提升复杂任务规划能力
One STEP at a time: Language Agents are Stepwise Planners
- 设计四模块协同框架,分步拆解任务并利用经验优化决策
- 在ScienceWorld上达成67.4分,18个任务中成功完成12个
- 适合需要长期规划的动态环境任务研究者参考
语言代理在动态环境中展现出了出色的适应性,能够执行复杂任务。然而,尽管大型语言模型蕴含丰富知识,这类代理在需要规划的任务上仍表现不足。本文提出STEP框架,通过高效学习过往经验来增强语言代理未来的规划能力。STEP由四个相互关联的组件构成:规划器负责分解任务并提供见解;执行器生成动作候选;评估器确保动作符合过往经验规则;记忆模块存储经验以指导未来决策。在ScienceWorld基准测试中,STEP持续优于现有最先进模型,整体得分达到67.4,成功完成18项任务中的12项。结果表明,STEP具备提升语言代理规划能力的潜力,为动态环境中更复杂的任务求解提供了新路径。
原文摘要 · Abstract (English)
Language agents have shown promising adaptability in dynamic environments to perform complex tasks. However, despite the versatile knowledge embedded in large language models, these agents still fall short when it comes to tasks that require planning. We introduce STEP, a novel framework designed to efficiently learn from previous experiences to enhance the planning capabilities of language agents in future steps. Concretely, STEP functions through four interconnected components. First, the Planner takes on the task, breaks it down into subtasks and provides relevant insights. Then the Executor generates action candidates, while the Evaluator ensures the actions align with learned rules from previous experiences. Lastly, Memory stores experiences to inform future decisions. In the ScienceWorld benchmark, our results show that STEP consistently outperforms state-of-the-art models, achieving an overall score of 67.4 and successfully completing 12 out of 18 tasks. These findings highlight STEP's potential as a framework for enhancing planning capabilities in language agents, paving the way for more sophisticated task-solving in dynamic environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。