模型在任务上默默学习关键中间步骤,即使损失不下降也已悄然进步。
Quiet Feature Learning in Algorithmic Tasks
- 通过探测内部表征发现:模型先学‘静默特征’,即不直接影响损失的中间计算。
- 在大量算力下损失几乎不变,随后突然下降,体现显著相变现象。
- 这些静默特征对任务表现有因果必要性,适合关注训练隐性进展的研究者。
我们在十个基础算法任务上训练基于Transformer的语言模型,观察到验证损失曲线出现显著偏离经典幂律缩放趋势的相变现象:在大范围算力下损失几乎不变,随后突然下降。探测模型内部表征发现,在任务损失下降前,已习得一系列‘静默特征’,它们代表不直接改善输出损失的中间算法计算。消融实验表明,单个静默特征对任务性能具有因果必要性。结果表明,可观的表征进展可能被看似平坦的损失曲线掩盖,挑战了以交叉熵作为学习代理的普遍做法,呼吁采用更丰富的诊断手段监控模型训练。
原文摘要 · Abstract (English)
We train Transformer-based language models on ten foundational algorithmic tasks and observe pronounced phase transitions in their loss curves that deviate from established power-law scaling trends. Over large ranges of compute, the validation loss barely improves, then abruptly decreases. Probing the models' internal representations reveals that quiet features are learned prior to any decrease in task loss. These quiet features represent intermediate algorithmic computations that do not by themselves improve the output loss. Ablation experiments demonstrate that individual quiet features are causally necessary for task performance. Our results demonstrate that substantial representational progress can remain hidden beneath an apparently flat loss curve, challenging the prevailing use of cross-entropy as a proxy for learning and motivating richer diagnostics for monitoring model training.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。