让小模型也能自我反思进化,提升任务表现
Improving Retrospective Language Agents via Joint Policy Gradient Optimization
- 双阶段联合优化,融合模仿学习与强化学习
- 在多个环境中显著提升任务完成率与决策质量
- 适合想降低大模型依赖、实现持续进化的研究者
近期研究中,大语言模型(LLMs)推动了自主智能体的发展。然而,当前基于提示的智能体高度依赖大规模闭源模型;尽管微调能提升小型模型能力,但其缺乏自我反思与进化潜力。为此,我们提出RetroAct框架,通过联合优化任务规划与自我反思演化能力,设计两阶段联合优化流程,结合模仿学习与强化学习,并提出一种带模仿学习正则化的离策略联合策略梯度算法,提升数据效率与训练稳定性。实验表明,RetroAct显著提升开源模型性能,减少对闭源模型依赖,使微调后的智能体具备持续学习与进化能力,在多个测试环境中均实现任务表现与决策过程的显著改善。
原文摘要 · Abstract (English)
In recent research advancements within the community, large language models (LLMs) have sparked great interest in creating autonomous agents. However, current prompt-based agents often heavily rely on large-scale LLMs. Meanwhile, although fine-tuning methods significantly enhance the capabilities of smaller LLMs, the fine-tuned agents often lack the potential for self-reflection and self-improvement. To address these challenges, we introduce a novel agent framework named RetroAct, which is a framework that jointly optimizes both task-planning and self-reflective evolution capabilities in language agents. Specifically, we develop a two-stage joint optimization process that integrates imitation learning and reinforcement learning, and design an off-policy joint policy gradient optimization algorithm with imitation learning regularization to enhance the data efficiency and training stability in agent tasks. RetroAct significantly improves the performance of open-source models, reduces dependency on closed-source LLMs, and enables fine-tuned agents to learn and evolve continuously. We conduct extensive experiments across various testing environments, demonstrating RetroAct has substantial improvements in task performance and decision-making processes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。