发现最优搜索框架需动态调整,而非固定模板。
Automated Discovery Has No Universally Superior Harness

- 拆解演化搜索组件,系统对比30种配置组合
- 无通用最优框架,多数表现不如简单方案
- 早期表现可预测最终结果,支持动态资源重分配
自主发现系统如OpenEvolve和TTT-Discover常被当作通用搜索框架使用。但这些系统是档案管理、父代选择、探索策略与预算分配等多种设计选择的混合体。由于实验成本高且具有随机性,现有对比通常依赖过少独立试验,难以区分方法改进与运行波动。本文将OpenEvolve式演化搜索与TTT-Discover框架分解为基本组件,在12个模型-问题对上系统评估30种预算匹配的框架,共执行超过310万次LLM推理,并采用重复试验统计分析。结果表明,发现框架存在泛化问题:无固定框架在所有场景下均最优,且多数OpenEvolve变体表现弱于简单替代方案。因此,框架选择应视为超参数,需针对具体任务和模型定制。此外,早期发现进展可有效预测最终性能,据此提出一种预算匹配的自适应分配实验:同时启动多个框架,淘汰表现差的初步运行,并将计算资源重新分配给表现更好的候选者,其效果优于随机固定框架或非自适应集成。研究建议从固定框架转向基于早期表现的在线自适应机制。所有运行数据集及基线零分布已公开,供未来框架评估复用。
原文摘要 · Abstract (English)
Autonomous discovery systems such as OpenEvolve and TTT-Discover are often used as general-purpose harnesses. However, in practice these are composite systems combining several design choices about archives, parent selection, exploration, and budget allocation into a single recipe. Because discovery runs are expensive and inherently stochastic, existing harnesses are often compared using too few independent trials to distinguish key methodological improvements from run-to-run variance. We systematically decompose OpenEvolve-style evolutionary search and the TTT-Discover search harness into its constituent components and systematically evaluate 30 budget-matched harnesses across 12 model-problem pairs using more than 3.1 million LLM rollouts and repeated-trial statistical analysis. Our results show that discovery harnesses have a generalization problem: No fixed harness is reliably superior across the evaluated model-problem pairs, and variants of OpenEvolve generally underperform simpler alternatives. Thus, harness choice is better viewed as a hyperparameter rather than as a universal recipe, and should be tailored to the specific problem and underlying model. We also find that early discovery progress predicts final performance, and use this property to present a budget-matched adaptive-allocation experiment that starts multiple harnesses, prunes weak partial runs, and reallocates compute to stronger survivors, outperforming both commitment to a randomly sampled fixed harness and a non-adaptive harness ensemble. Together, these results motivate shifting from fixed harness selection to online adaptation guided by early performance. We release all run pools including baseline null distributions for every model-problem pair as reusable statistical infrastructure against for future harness proposals.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。