arXiv:2606.11387cs.CLcs.AI2026-06

通过分阶段筛选,用低成本实验高效找到稳定最优模型配置。

Small Experiments, Cheaper Decisions: A Case Study in Staged Promotion for Micro-Pretraining

论文配图:Small Experiments, Cheaper Decisions: A Case Study in Staged Promotion for Micro-Pretraining
图 1 · 摘自论文原文
  • 设计多阶段预算筛选流程,逐步淘汰表现不稳定的配置。
  • 12小时最终验证中,指定配置在所有条件下均排名第一,且符合严格规则。
  • 适合追求低成本高效实验的模型训练团队,尤其关注可复现性与成本控制。

短时预训练可降低实验成本,但也可能过度青睐仅在小预算下表现好的配置。本文研究了一种可审计的分阶段晋升协议,在两个异构主机(Windows A100 和 Linux L40S)上对固定微预训练运行器进行评估。从12个预先筛选的配置出发,采用2分钟、5分钟、10分钟、60分钟和12小时的分阶段预算,且在昂贵延续前冻结晋升规则。早期阶段(5分钟和10分钟)排名受主机影响,12小时最终排名并非10分钟阶段的平均最佳配置。因各阶段种子范围不同,这些变化为实际晋升提供了操作证据,而非单一种子下的曲线变化。60分钟阶段的桥接配置在所有四个主机-种子组合中均排名第一;12小时最终确认中,该配置仍居首位,而贪婪对比项未满足0.010的验证比特率近似等价规则,且更便宜的d8/ar48哨兵也未通过0.020的平均差距规则。完整流程共消耗169.2训练GPU小时,12小时分支为144 GPU小时。若继续全部4个60分钟候选者将耗192 GPU小时,全部9个10分钟候选者则需432 GPU小时——后者为未执行的反事实计算,不代表被跳过的配置无法超越参考项。结果为可控成本分配的发现,非全局最优或自适应超参优化方法的性能宣称。

原文摘要 · Abstract (English)

Short pretraining runs can reduce experimental cost, but they can also over-promote configurations that only look strong at tiny budgets. We study an auditable staged-promotion protocol for a fixed micro-pretraining runner on two heterogeneous host blocks: Windows A100 and Linux L40S. Starting from twelve prior-screened configurations, we use staged budgets of 2 minutes, 5 minutes, 10 minutes, 60 minutes, and 12 hours, with frozen promotion rules before expensive continuations. The early screens are intentionally treated as unstable: the 5- and 10-minute rankings are host-sensitive, and the eventual 12-hour top-ranked condition is not the mean-best condition at the replicated 10-minute gate. Because seed ranges differ across stages, these changes are operational promotion evidence, not within-seed curves. A replicated 60-minute gate keeps the Staged Factorial Screening bridge reference in the promoted set, where it ranks first in all four 60-minute host-seed cells. In the final 12-hour confirmation package, the bridge condition ranks first in all four host-seed cells across two seeds; the greedy comparator does not meet the frozen 0.010 val_bpb near-equivalence rule; and the cheaper d8/ar48 (depth-8, aspect-48) sentinel does not meet the frozen 0.020 mean-gap rule. The executed 12-hour branch spends 144 GPU-hours, and the full staged protocol records 169.2 training GPU-hours including screening stages. Continuing all four 60-minute candidates would spend 192 GPU-hours, while continuing all nine replicated 10-minute candidates would spend 432 GPU-hours. The latter numbers are accounting counterfactuals for unrun continuations, not evidence that skipped candidates could not have overtaken the reference. The result is a bounded cost-allocation finding, not a claim of global optimality, capacity-normalized superiority, or superiority over adaptive hyperparameter optimization methods.

模型训练实验优化成本控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。