arXiv:2503.02682cs.CLcs.AI2025-03EMNLP被引 42

用元计划优化提升大模型智能体的规划能力,避免幻觉且无需重训练。

MPO: Boosting LLM Agents with Meta Plan Optimization

  • 通过元计划提供高层指导,实现规划能力持续优化。
  • 在两个任务上显著超越现有基线,提升任务完成效率与泛化能力。
  • 即插即用,适用于新场景,减少人工干预和重训练成本。

大语言模型(LLM)的发展使基于LLM的智能体能够成功应对交互式规划任务。然而,现有方法常出现规划幻觉,且需为每个新智能体重新训练。为此,我们提出元计划优化(Meta Plan Optimization, MPO)框架,通过直接引入显式指导来增强智能体的规划能力。与依赖复杂知识、需大量人工或质量难保障的方法不同,MPO利用高层通用指导(元计划)辅助规划,并根据智能体任务执行反馈持续优化元计划。在两个代表性任务上的实验表明,MPO显著优于现有基线。分析显示,MPO提供即插即用方案,有效提升任务完成效率与未见场景下的泛化能力。

原文摘要 · Abstract (English)

Recent advancements in large language models (LLMs) have enabled LLM-based agents to successfully tackle interactive planning tasks. However, despite their successes, existing approaches often suffer from planning hallucinations and require retraining for each new agent. To address these challenges, we propose the Meta Plan Optimization (MPO) framework, , which enhances agent planning capabilities by directly incorporating explicit guidance. Unlike previous methods that rely on complex knowledge, which either require significant human effort or lack quality assurance, MPO leverages high-level general guidance through meta plans to assist agent planning and enables continuous optimization of the meta plans based on feedback from the agent's task execution. Our experiments conducted on two representative tasks demonstrate that MPO significantly outperforms existing baselines. Moreover, our analysis indicates that MPO provides a plug-and-play solution that enhances both task completion efficiency and generalization capabilities in previous unseen scenarios.

大模型智能体规划优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。