arXiv:2604.04281cs.AI2026-04被引 1

宽模型初始化选不好,效果反而更差。

Preservation Is Not Enough for Width Growth: Regime-Sensitive Selection of Dense LM Warm Starts

论文配图:Preservation Is Not Enough for Width Growth: Regime-Sensitive Selection of Dense LM Warm Starts
图 1 · 摘自论文原文
  • 对比多种初始化方式,发现直接复制最有效
  • 长序列确定性生成中,非克隆结构表现更好
  • 随机生成场景下,早期脱离原模型反而会拖后腿

宽度扩展为重用小型因果语言模型检查点提供了一条实用路径,但仅靠零步保全(zero-step preservation)无法解决宽化初始值的选择问题。本文将密集宽度增长视为对完整训练状态(包括复制权重、优化器动量和调度器状态)的候选选择问题。在小规模TinyStories代理实验中,比较了精确复制、微扰、非对称重置和结构化非克隆四种初始化方式,在相同延续预算下评估零步保全、短延迟探测指标与下游延续效用。结果表明:在16步探测中,精确复制对称初始化在所有任务中排名第一;在种子0的1000和2000步及种子1的2000步处,其随机延续表现优异。然而,在确定性128步延续中,结构化非克隆方案胜出。因此,早期脱离继承的克隆子空间并非普适选择标准:它在长确定性延续中有益,但在短延迟和随机延续中误导。结论是:在此尺度下,单纯保全是不够的,最佳替代信号取决于任务类型和延迟预算。

原文摘要 · Abstract (English)

Width expansion offers a practical route to reuse smaller causal-language-model checkpoints, but selecting a widened warm start is not solved by zero-step preservation alone. We study dense width growth as a candidate-selection problem over full training states, including copied weights, optimizer moments, and scheduler state. In a small-scale TinyStories proxy, we compare exact-copy, perturbative, asymmetric-reset, and structured non-clone warm starts under matched continuation budgets. We evaluate zero-step preservation, short-lag probe metrics, and downstream continuation utility in deterministic and stochastic regimes. The picture is mixed and partially replicated through a reduced-pool seed-1 check. Exact-copy symmetric warm starts rank first in every completed 16-step probe and in the completed stochastic 128-step continuations at seed-0 steps 1000 and 2000 plus reduced seed-1 step 2000. By contrast, the structured non-clone challenger wins deterministic 128-step continuation. Early escape from the inherited cloned subspace is therefore not a universal selector: it helps in long deterministic continuation, but it misleads at short lag and under stochastic continuation. The result is narrow but useful: for dense width growth at this scale, preservation is not a universal ranking criterion, and the best replacement signal depends on both regime and lag budget.

模型扩展初始化策略语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。