用专业智能体自动优化训练配方,无需人工干预即可提升模型性能。
Auto Research with Specialist Agents Develops Effective and Non-Trivial Training Recipes

- 通过专业化智能体分工协作,基于实验反馈迭代优化训练配方。
- 在3个任务中分别降低0.81%、提升38.7%、减少4.59%,均经外部评估验证。
- 全程自动执行代码提交与修正,适合自动化机器学习研究场景。
我们研究以外部测量为驱动的闭环自动研究流程。每次实验提交包含假设、可执行代码修改、评估器拥有的结果及反馈,用于指导下一提案。输出不是生成论文或单一模型检查点,而是可审计的提案、代码差异、实验记录、评分和失败标签轨迹。我们用专业智能体实现该循环,划分配方空间并共享测量传承。核心发现是,传承反馈使智能体将评估结果(如崩溃、预算超支、规模失败、准确率未达标)转化为后续程序级配方修改,而非一次性建议。在完成一次设置后,经过1,197次主实验运行及600次Parameter Golf对照实验,人类未参与提案选择、配方编辑、分数修改或失败修复。在三个主实验中,同一实验循环分别将Parameter Golf验证比特率降低0.81%,使NanoChat-D12 CORE提升38.7%,将CIFAR-10 Airbench96运行时间减少4.59%,各任务由独立评估器与合法性检查验证。完整轨迹包含对157次主实验提交的严格架构域审计,以及诸如NanoChat注意力核路径更改等程序重写。在此范围内,系统自主编写代码、提交实验、吸收反馈、整合已知技术并改进公开初始配方。
原文摘要 · Abstract (English)
We study auto research as a closed empirical loop driven by external measurement. Each submitted trial carries a hypothesis, an executable code edit, an evaluator-owned outcome, and feedback that shapes the next proposal. The output is not a generated paper or a single model checkpoint, but an auditable trajectory of proposals, code diffs, experiments, scores, and failure labels. We instantiate this loop with specialist agents that partition recipe surfaces and share measured lineage across trials. The central empirical finding is that lineage feedback lets agents turn evaluator outcomes, including crashes, budget overruns, size failures, and accuracy-gate misses, into later program-level recipe edits rather than one-shot suggestions. Across 1,197 headline-run trials plus 600 Parameter Golf control trials after one-time setup and launch, humans did not choose proposals, edit recipes, override scores, or repair failed trials during the search. In the three headline runs, the same submitted-trial loop reduces Parameter Golf validation bpb by $0.81\%$, raises NanoChat-D12 CORE by $38.7\%$, and reduces CIFAR-10 Airbench96 wallclock by $4.59\%$, with each task measured by its own external evaluator and legality checks. The trace includes a strict architecture-domain audit of 157 headline-run submissions and program rewrites such as a NanoChat attention-kernel path change. Within this scope the loop autonomously writes code, submits experiments, absorbs feedback, applies and combines known techniques inside each environment, and improves public starting recipes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。