在有限算力下,研究小模型训练过程中的波动与退化现象。
A Quantitative Experimental Repeated Measures Study of Training Dynamics in a Small Llama Style Language Model Under a Compute-Aware Token Budget

- 采用重复测量设计,分阶段追踪损失与困惑度变化
- 早期快速下降后出现非单调上升,最终损失升至3.9010
- 适合关注训练稳定性与算力效率的研究者
本研究在固定算力约束的令牌预算下,分析了一个小型Llama风格语言模型的训练动态。通过六次独立训练运行,在426万参数模型上使用TinyStories语料库进行基于CPU的全精度训练,目标总训练令牌数约2000万。在21个令牌区间内采集数据,共获得126个种子-区间观测值。重复测量方差分析显示验证损失、验证困惑度和滚动波动性在区间间存在显著差异。描述性轨迹表明,训练初期迅速提升,但后期出现非单调下降;平均验证损失从初始的8.3552降至约400万令牌时的2.7996,但在最终检查点回升至3.9010。验证困惑度呈现相同趋势。进一步分析发现频繁出现验证损失回撤,且无稳定阶段证据。结果表明,在算力受限环境下,仅看终点性能会掩盖训练过程中的不稳定性、退化与收益递减。建议结合区间级遥测数据评估模型训练状态。
原文摘要 · Abstract (English)
This study examines training dynamics in a small Llama-style language model trained under a fixed, compute-constrained token budget. Rather than evaluating efficiency solely through endpoint performance, the study uses a quantitative experimental repeated measures design to analyze how validation loss, validation perplexity, rolling volatility, backslide behavior, spike behavior, and between-seed variability change across token-based training intervals. Six independent training runs were conducted on a 4.26-million-parameter model using the TinyStories corpus, CPU-based full-precision training, and a target budget of approximately 20 million cumulative training tokens. Metrics were collected across 21 intervals, producing 126 seed-by-interval observations. Repeated measures ANOVA showed statistically significant interval effects for validation loss, validation perplexity, and rolling volatility. Descriptive trajectories revealed rapid early improvement followed by non-monotonic degradation during later training intervals. Mean validation loss decreased from 8.3552 at initialization to 2.7996 near 4 million tokens, but increased to 3.9010 by the final checkpoint. Validation perplexity followed the same pattern, falling sharply early in training before rising later. Derived telemetry further showed recurrent validation-loss backslides and no interval-summary evidence of a stable phase under the predefined criteria. These findings suggest that compute-aware language model evaluation should examine training trajectories rather than endpoint metrics alone. In constrained compute settings, additional token exposure may increase computational cost without producing proportional generalization gains, and interval-level telemetry can reveal instability, regression, and diminishing returns that final metrics may obscure.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。