arXiv:2506.15872cs.LG2025-06被引 15

通过分解损失变化,发现语言模型训练中的隐藏能力跃迁。

Hidden Breakthroughs in Language Model Training

  • 用POLCA方法在低秩训练子空间分解损失变化
  • 识别出训练中多个可解释的能力跃迁簇
  • 适合研究模型学习动态与无监督可解释性

训练过程中损失曲线通常平滑,但突现的不连续点可能代表关键概念突破。本文指出,类似突破在训练中频繁发生,却被单一标量损失掩盖。为此提出POLCA方法,将损失变化沿低秩训练子空间的任意基进行分解,识别出具有相似损失变化模式的样本簇,将整体损失拆解为若干概念相似数据组的损失。在合成算术和自然语言任务上验证,POLCA能恢复出可解释的能力跃迁簇,证明这些隐藏相变对无监督可解释性具有潜力。

原文摘要 · Abstract (English)

Loss curves are smooth during most of model training, so visible discontinuities stand out as possible conceptual breakthroughs. Studying these breakthroughs enables a deeper understanding of learning dynamics, but only when they are properly identified. This paper argues that similar breakthroughs occur frequently throughout training but they are obscured by a loss metric that collapses all variation into a single scalar. To find these hidden transitions, we introduce POLCA, a method for decomposing changes in loss along arbitrary bases of the low-rank training subspace. We use our method to identify clusters of samples that share similar changes in loss during training, disaggregating the overall loss into that of smaller groups of conceptually similar data. We validate our method on synthetic arithmetic and natural language tasks, showing that POLCA recovers clusters that represent interpretable breakthroughs in the model's capabilities. We demonstrate the promise of these hidden phase transitions as a tool for unsupervised interpretability.

模型训练损失分析可解释性相变

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。