arXiv:2510.05608cs.CL2025-10ACL被引 18

让大模型有规划地完成复杂任务,提升效率与准确性。

A Goal Without a Plan Is Just a Wish: Efficient and Effective Global Planner Training for Long-Horizon Agent Tasks

  • 用自动生成的高质量计划训练全局规划器,无需人工标注。
  • 在三项长程任务中表现超越现有方法,训练成本降低8倍。
  • 适合需要长期规划的智能体系统,如自动化流程设计。

基于大语言模型(LLM)的智能体在长程任务中因缺乏全局规划能力,常陷入盲目试错和生成幻觉动作。本文提出一种“计划-执行”框架及EAGLET训练方法,通过两步法构建可即插即用的全局规划器:首先利用同源共识过滤策略从先进LLM生成高质量计划,并以微调作为冷启动;随后采用基于规则的强化学习,引入执行器能力提升奖励,使规划器能应对不同难度的任务指令。在三个长程任务上的实验表明,配备该规划器的执行器显著优于现有方法,达到新基准性能。同时,EAGLET相较基于RL的基线将训练成本降低8倍,且无需人工干预或额外数据,提供高效有效的解决方案。

原文摘要 · Abstract (English)

Agents based on large language models (LLMs) struggle with brainless trial-and-error and generating hallucinatory actions due to a lack of global planning in long-horizon tasks. In this paper, we introduce a plan-and-execute framework and propose EAGLET, an efficient and effective planner training method to enhance the executor agent's planning abilities without human effort. Specifically, we train a plug-and-play global planner through a two-step process: we first synthesize high-quality plans from an advanced LLM using our proposed homologous consensus filtering strategy, and apply fine-tuning as a cold start. Moreover, we further improve the planner with a rule-based reinforcement learning stage using a novel executor capability gain reward, ensuring it can handle task instructions of varying difficulty. Experiments on three long-horizon agent tasks show that executor agents equipped with our planner outperform existing methods, achieving new state-of-the-art performance. Meanwhile, EAGLET reduces training costs by 8x compared to RL-based baselines, and it does not require manual effort or extra training data, offering an efficient and effective solution.

大模型规划生成强化学习长程任务

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。