30B大模型自主训练,自动优化且发现评估指标失效
A-Evolve-Training: Autonomous Post-Training of a 30B Model

- 全程无人干预,自动完成数据、策略、评估闭环迭代
- 在推理挑战赛中达0.86分,接近人类最优的0.87分
- 能识别评估指标失真并主动调整策略,展现自主发现能力
后训练前沿模型通常需数周人工操作:设计数据与配方、启动实验、分析评估结果、决定保留内容。我们报告一个完全自主的系统,在多周内对30亿参数的Nemotron模型完成四轮后训练,无需人工介入。该系统生成的模型在公开的NVIDIA Nemotron-Reasoning Challenge排行榜上取得0.86的保留分数,仅次于人类最优的0.87,位列约4000个提交中第8名。更关键的是,系统发现其内部开发指标在最弱领域已不再反映外部性能——尽管候选方案使开发指标创纪录提升,但外部目标未变。于是系统自行修正搜索策略,不再最大化开发指标,转而寻找能降低误导性代理指标同时提升外部目标的干预措施。这被视为直接、可审计的证据:规模化自主循环不仅能优化,还能实现发现——它察觉到测量框架已失真,并改变何为有效证据的标准。我们认为,真正具备‘递归自我改进’能力的系统,必须最终能端到端完成前沿模型的后训练;本工作是达成这一标准的一个实证。我们不声称首次实现与人类研究者的自主匹敌。我们的主张更具体且可验证:据我们所知,这是首个公开报道的在该规模(30B)上实现的自主后训练,此前公开的自主机器学习研究均限于类似GPT-2(约124M)的预算。同一系统也成功后训练了120B和550B的Nemotron;因无公开人类基准,仅证明系统在此规模仍能闭环运行,有效性待未来人类基准出现后再判断。
原文摘要 · Abstract (English)
Post-training a frontier model is normally weeks of human work: proposing data and recipe changes, launching runs, reading evals, deciding what to keep. We report an autonomous system that runs this loop with no human in the loop, post-training a 30B Nemotron across four rounds over multiple weeks. The autonomously produced model reaches a held-out score of 0.86 against the top human submission's 0.87 on the public NVIDIA Nemotron-Reasoning Challenge leaderboard, placing 8th of ~4000 at the time of writing. More striking than the number: the loop detected that its own dev metric had stopped tracking external performance on the weakest domain -- candidates drove dev to record highs without moving the external target -- and revised its own search policy, no longer maximizing dev but seeking interventions that lowered the now-misleading proxy while improving the external target. We treat this as direct, auditable evidence that a scaled autonomous loop can produce discovery, not only optimization: it detected that its measurement frame had become misleading and changed what counted as evidence. We take the operational view that any system worth the "recursive self-improvement" label must eventually perform end-to-end post-training of a frontier-class model; this is one datapoint of that bar being cleared. We do not claim a "first autonomous match" of human researchers. The claim we make is narrower and auditable: to our knowledge, this is the first publicly reported autonomous post-training run at this scale, where prior public autonomous-ML-research demonstrations sit at GPT-2-class (~124M) budgets. The same system also post-trains the 120B and 550B Nemotron; with no public human baseline there, this shows only that the loop closes at that scale, not that its output is competitive -- infrastructure evidence, with the effectiveness claim deferred until a comparable human anchor exists.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。