小模型指导大模型分阶段突破技能瓶颈,大幅降低训练成本。
Small Models Scout Bottleneck Order for Large-Model Data Control

- 用小模型追踪大模型训练瓶颈顺序,动态引导优化路径。
- 在Qwen2.5-1.5B上平均节省56.2%训练token,70M→12B迁移中节省39.4%。
- 适合需要高效数据控制的大型模型训练场景,尤其关注技能渐进提升。
小模型代理常用于为大规模训练识别数据混合方案。本文探究其训练轨迹是否揭示另一可迁移结构:大模型应如何按顺序解决技能瓶颈。我们提出首达技能训练(first-passage skill training),每个监控技能设定目标下限,目标是最小化达到所有下限所需的训练tokens。引入LogFloor闭环控制器,每轮聚焦当前瓶颈,生成有序的瓶颈解决轨迹。在五个bAbI技能切片上,LogFloor使Qwen2.5-1.5B平均减少56.2%的训练开销。在70M到12B模型迁移中,仅三轮重播70M小模型路径即可在全部八次实验中达成目标,相较配对均值节省30.9%,联合训练tokens减少39.4%,源成本核算下节省37.6%。在MMLU-control任务中,冻结的小模型路径在所有八次12B运行中均成功。将路径压缩为静态边际混合或反转阶段顺序会丧失大部分优势,而仅使用瓶颈标签仍具部分效用。结果表明,阶段有序的瓶颈解决是一种可迁移的课程结构,适用于监控技能目标的训练。
原文摘要 · Abstract (English)
Small proxy models are commonly used to identify data mixtures for larger-scale training. We ask whether their training trajectories reveal another transferable structure: the order in which larger models should resolve skill bottlenecks. We formulate first-passage skill training, where each monitored skill has a target floor and the objective is to minimize the tokens required to reach all floors. We introduce LogFloor, a closed-loop controller that directs each round toward current bottlenecks, producing phase-ordered resolution trajectories. Across five bAbI skill slices on Qwen2.5-1.5B, LogFloor reduces token cost by 56.2% on average. In 70M-to-12B transfer, three-round replay of a 70M scout path reaches every floor in all eight target runs, saving 30.9% by pair mean, 39.4% in pooled training tokens, and 37.6% under source-cost accounting. On MMLU-control, a frozen scout path succeeds across all eight 12B runs. Collapsing a path to its static marginal mixture or reversing its phase order removes most benefits, while bottleneck labels alone remain partially useful. These results identify phase-ordered bottleneck resolution as a transferable curriculum structure for monitored skill-targeted training.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。