arXiv:2606.07856cs.LG2026-06被引 1

自训练让模型更擅长表达已有能力,而非突破能力边界。

Teacher-Free Self-Training Amplifies but Does Not Compound: A Pass@$K$ Crossover on a Free-Verifier Domain

  • 用生成器、判别器和免费验证器构成无教师自训练系统。
  • 自训练提升性能上限但不加速收敛,且在大预算下基础模型反而更优。
  • 证明是能力放大而非能力叠加,适合研究模型泛化与自训练机制。

当语言模型基于自身验证输出进行训练时,其能力是否超越基础模型,还是仅更高效地表达原有能力?我们通过一个无教师的‘星群’系统(生成器、学习型判别器、免费精确验证器)在类似FlashFill的‘陷阱门’领域中解答此问题,该领域可低成本合成验证过的(问题, 解答)对,难以逆向,且可精确验证。所有实验在单张24GB显卡上运行,仅使用4比特的Qwen3-4B模型,无更大模型参与循环。我们报告三个发现:(i) 判别器引导选择比验证器过滤的best-of-$k$提升9.1个百分点(6/6种子),全部增益集中于候选答案在保留输入上存在分歧的任务;(ii) 每轮自训练提升性能上限但不加速,增益随剩余潜力下降而减缓,且在$K=4$独立训练轨迹中表现一致;(iii) 该领域无清晰零能力边界,故传统‘0%→上升=涌现’测试无效。通过测量pass@$K$交叉点得出结论:训练模型在操作预算(pass@$8$)下胜出,但基础模型在大预算(pass@$64$)下始终反超,说明自训练仅集中概率质量,而非拓展能力范围。这是放大,而非复合。($K=4$为示意性结果,非跨轨迹稳健置信区间。)

原文摘要 · Abstract (English)

When a language model trains on its own verified outputs, does it acquire capability beyond its base, or merely get better at expressing capability the base already had? We make the question decidable with a teacher-free "constellation" -- a generator, a learned critic, and a free exact verifier -- on a FlashFill-style "trapdoor" DSL, where verified (problem, solution) pairs are cheap to synthesize, hard to invert, and free to check exactly. Everything runs on one 4-bit Qwen3-4B on a single 24 GB GPU, with no model in the loop larger than the base. We report three findings. (i) Critic-guided selection beats verifier-filtered best-of-$k$ by $+9.1$ pp ($6/6$ seeds), with the entire gain localized to tasks where candidates disagree on held-out inputs. (ii) Per-round STaR self-training raises the ceiling but never accelerates -- the gain tracks remaining headroom and decelerates across $K=4$ independent training trajectories. (iii) The domain has no clean zero-capability frontier, so the usual "$0\% \to$ climb $=$ emergence" test is invalid here. A measured pass@$K$ crossover settles the diagnosis: the trained model wins at the operating budget (pass@$8$) but the base overtakes it at a large budget (pass@$64$) on every trajectory, so self-training concentrates probability mass rather than expanding reach. This is amplification, not compounding. ($K=4$ is indicative, not yet a robust across-trajectory CI.)

自训练能力放大验证机制语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。