arXiv:2608.09629cs.AI2026-08

用大模型自主设计优化流程,比固定步骤更高效且省资源。

Rethinking Self-Evolving Agents: Do We Still Need Prescribed Optimization Pipelines?

论文配图:Rethinking Self-Evolving Agents: Do We Still Need Prescribed Optimization Pipelines?
图 1 · 摘自论文原文
  • 让大模型自动生成优化路径,不再依赖预设步骤。
  • 在14组对比中赢12次,仅用34.3%的交互预算。
  • 适合能力强的模型,对弱模型效果有限。

自演化代理通常依赖预设的优化流程:框架决定如何收集证据、修改持久性成果、选择候选方案并终止。我们探讨当前沿模型作为优化器时,这种任务特定流程是否仍必要。为此提出开放式优化(OEO),固定目标、交互权限、资源预算、数据边界和评估方式,允许优化器在线组合改进过程。在8个基准-目标模型设置下进行14组头对头比较,由GPT-5.5驱动的OEO取得12胜1平1负(差距0.21个百分点)的成绩,仅使用技能优化(SkillOpt)34.3%的交互令牌预算。单次无交互对照实验表明,性能提升并非源于单一先验重写。然而,委托存在能力边界:中等能力优化器下,SkillOpt表现更优;弱优化器无法通过不变的OEO接口运行。完全可测的OEO-SkillOpt配对分析显示,预设流程更多影响优化轨迹的一致性,而非最终行为。研究重新定义预设流程为能力依赖的支撑结构:外部约束仍必要,但足够强大的优化器可从可度量反馈自主构建持续改进路径。

原文摘要 · Abstract (English)

Self-evolving agents are usually built around prescribed optimization pipelines: the framework decides how to gather evidence, revise a persistent artifact, select candidates, and stop. We ask whether this task-specific procedure remains necessary when a frontier model acts as the optimizer. We introduce Open-Ended Optimization (OEO), which keeps the objective, permitted interactions, resource budget, data boundary, and evaluation fixed while allowing the optimizer to compose the improvement process online. We compare OEO with two complementary prescribed approaches: SkillOpt, a staged pipeline with bounded edits, and GEPA, a reflective evolutionary search. Across 14 head-to-head comparisons over 8 benchmark-target-model settings, GPT-5.5-driven OEO records 12 wins, 1 tie, and 1 narrow loss of 0.21 percentage points. It uses a median 34.3 percent of SkillOpt's configured target-interaction token budget. A one-shot, zero-interaction control shows that the gains are not explained by a single prior-driven rewrite. However, delegation has a capability boundary: SkillOpt outperforms OEO with a medium optimizer, and a weak optimizer cannot operate through the unchanged OEO interface. In the fully instrumented OEO-SkillOpt pair, trajectory analysis further shows that prescription changes how optimization proceeds more consistently than it changes final behavior. Together, these findings recast prescribed pipelines as capability-dependent scaffolding: essential constraints remain external, but a sufficiently capable optimizer can compose the route from measurable feedback to persistent improvement.

自演化优化流程大模型零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。