arXiv:2504.04785cs.AI2025-04被引 19

用小模型自动设计工作流,让大模型表现更优

Weak-for-Strong: Training Weak Meta-Agent to Harness Strong Executors

  • 用强化学习训练小模型自动设计任务流程
  • 7B小模型仅用1小时训练,性能超基线2.9%~24.6%
  • 无需人工干预,通用性强,适合高效利用大模型

高效利用现代大语言模型(LLMs)的能力日益困难,尤其当直接微调成本高昂且不切实际时。现有无训练方法(如手动或自动化设计的工作流)通常需要大量人力或效果不佳。本文提出弱模型赋能强执行者框架(W4S),通过小型低成本语言模型设计并优化用于调用更强模型的工作流。W4S将工作流设计建模为多轮马尔可夫决策过程,并引入强化学习进行代理式工作流优化(RLAO),训练一个弱元代理。通过与环境的迭代交互,元代理无需人工干预即可学习设计更高效的工作流。实验表明,仅用1小时在单张GPU上训练的7B元代理,在11个基准测试中表现优于最强基线2.9%~24.6%,成功提升GPT-3.5-Turbo和GPT-4o等前沿模型性能。值得注意的是,W4S在已见与未见任务上均表现出强泛化能力,提供了一种高效、高性能的替代直接微调强模型的方案。

原文摘要 · Abstract (English)

Efficiently leveraging of the capabilities of contemporary large language models (LLMs) is increasingly challenging, particularly when direct fine-tuning is expensive and often impractical. Existing training-free methods, including manually or automated designed workflows, typically demand substantial human effort or yield suboptimal results. This paper proposes Weak-for-Strong Harnessing (W4S), a novel framework that customizes smaller, cost-efficient language models to design and optimize workflows for harnessing stronger models. W4S formulates workflow design as a multi-turn markov decision process and introduces reinforcement learning for agentic workflow optimization (RLAO) to train a weak meta-agent. Through iterative interaction with the environment, the meta-agent learns to design increasingly effective workflows without manual intervention. Empirical results demonstrate the superiority of W4S that our 7B meta-agent, trained with just one GPU hour, outperforms the strongest baseline by 2.9% ~ 24.6% across eleven benchmarks, successfully elevating the performance of state-of-the-art models such as GPT-3.5-Turbo and GPT-4o. Notably, W4S exhibits strong generalization capabilities across both seen and unseen tasks, offering an efficient, high-performing alternative to directly fine-tuning strong models.

大模型调度强化学习工作流优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。