arXiv:2606.05186cs.LGcs.CL2026-06

通过分阶段实验筛选微预训练配方,节省算力并找到高效方案。

Staged Factorial Screening for Budget-Constrained Micro-Pretraining

  • 采用分阶段因子实验设计,快速识别关键影响因素。
  • 5分钟和10分钟内稳定识别出4个有效配方(D、A、B、C)。
  • 适合资源受限场景下的模型配方快速筛选,尤其在单卡环境下。

在预算受限的微预训练中,需在共享加速器上对大量候选配方进行初步筛选。本文研究分阶段分数因子工作流能否在短预算下识别出稳定的早期效应结构。在固定单卡训练循环下,共执行613次实验,涵盖2、5、10分钟的试运行与后续筛查,以及5和10分钟的16条件种子重跑、目标锚点检查、同主机贪婪与匹配成本随机基线、60分钟桥接包,及在双主机(Windows A100与Linux L40S)上持续24小时的有限锚点延续。总批次、深度和宽度的惩罚在短预算时最大,随预算增加逐渐缓解。在预设种子全屏族中,经组内贝尼哈伯校正后,D、A、B、C在5分钟和10分钟仍具非零估计值,而E不显著。随机搜索虽可命中强基准,但反复集中在低惩罚区域且无法归因因子。60分钟桥接锚点均值最低,但其优势受大模型容量影响。在12小时与24小时三锚点延续中,桥接方案样本均值最低,而非桥接顺序则仍受硬件影响。因此提出:用短周期设计实验识别高惩罚方向,重复验证优选锚点,并在缩小空间内局部优化。证据支持在两主机上24小时内以桥接为中心的推荐策略,而非硬件无关排序或通用超参数优化优势。

原文摘要 · Abstract (English)

Budget-constrained micro-pretraining often requires triaging many candidate recipes on a shared accelerator before larger search budgets are spent. We study whether a staged fractional-factorial workflow can recover stable early effect structure in this setting. On a fixed autoresearch-derived single-GPU training loop, we run 613 experiments across pilot and follow-up screens at 2, 5, and 10 minutes; full 16-condition seeded reruns at 5 and 10 minutes; targeted seeded anchor checks; same-host greedy and matched-cost random baselines; a 60-minute bridge package; and bounded Windows A100 and Linux L40S anchor continuations through 24 hours. Main penalties from total batch, depth, and width are largest at short budgets and relax as budget increases. Within the predeclared seeded full-screen families, D, A, B, and C retain non-zero estimates at 5 and 10 minutes after within-budget Benjamini-Hochberg correction, while E does not. Random search can reach strong incumbents in this 32-condition space, but repeatedly in the same low-penalty region and without factor attribution. The 60-minute bridge anchor has the lowest mean, although that package does not separate workflow refinement from the larger bridge model's capacity advantage. In bounded 12-hour and 24-hour three-anchor continuations on both hosts, the bridge has the lowest sample mean while the non-bridge ordering stays host-sensitive. We therefore present a bounded methods result: use short designed screens to identify high-penalty directions, confirm promising anchors under repeated runs, and refine locally inside the reduced space. The evidence supports a bridge-centered recommendation through 24 hours on two hosts, not hardware-invariant ranking or general hyperparameter-optimization superiority.

微预训练实验设计算力优化因子筛选

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。