基于群体比较的连续优化方法,提升智能体工作流自动构建效率
SPOGW: a Score-based Preference Optimization method via Group-Wise comparison for workflows
- 采用群体对比与连续空间优化,避免离散策略限制
- 在5个基准数据集上超越或媲美当前最先进方法
- 适合需要高效自动化工作流生成的研究者
大语言模型在多个领域展现出解决复杂问题的能力,常通过遵循结构化指令和多步骤流程的智能体工作流实现。然而,设计此类工作流需大量人工干预,制约其可扩展性和通用性。近期研究致力于减少人工参与,推动智能体工作流的自动化优化。但现有方法受限于表征能力弱、适应性差、可扩展性不足以及依赖成对比较范式,根源在于对离散优化技术的依赖。为此,本文提出一种基于得分的群体比较偏好优化方法(SPOGW),直接在基数奖励信号上操作,通过群体比较实现连续空间中的高效稳定优化。SPOGW结合迭代离线GRPO(ioGRPO)与优势掩码KL散度(mKL),通过强调策略响应的优势区域来调控训练更新。在涵盖数学推理、编程和问答的5个基准数据集上,SPOGW的表现匹配或超过当前最先进方法,为智能体工作流的自动构建与优化提供了可行且具有前景的新路径。
原文摘要 · Abstract (English)
Large language models (LLMs) have exhibited significant capabilities in addressing challenging problems throughout various fields, often through the use of agentic workflows that adhere to structured instructions and multi-step procedures. However, designing such workflows demands substantial manual effort, posing challenges to scalability and generalizability. Recent studies have aimed to minimize the human intervention needed for their construction, leading to advances in automated techniques for optimizing agentic workflows. However, current approaches are often constrained by their limited representational capacity, insufficient adaptability, weak scalability, and pairwise comparison paradigm -- issues that stem primarily from a dependence on discrete optimization techniques. To overcome these limitations, we introduce a new score-based preference approach, refereed as SPOGW, which operates directly on cardinal reward signals through group-wise comparison and enables more efficient and stable optimization in a continuous space. SPOGW incorporates Iterative offline GRPO (ioGRPO) with advantage-masked KL divergence (mKL), which regulates training update by placing greater emphasis on the advantageous regions of the policy response. In five benchmark datasets covering mathematical reasoning, coding, and question answering, SPOGW matches or exceeds the performance of current state-of-the-art approaches, presenting a viable and forward-looking methodology for automated generation and optimization of agentic workflows.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。