arXiv:2502.17356cs.LG2025-02被引 5

模型能力突现是随机性导致的分布变化,非规模阈值效应

Random Scaling of Emergent Capabilities

  • 用随机种子差异揭示能力突现本质是分布连续变化
  • 不同种子下同一任务可呈现平滑或突现两种趋势
  • 适合关注模型可复现性与评估可靠性的研究者

语言模型通常遵循平滑的缩放规律,但某些特定能力会出现性能突变。支持‘突现’观点者认为这些能力在特定规模被解锁,而另一些人则将其归因于表面指标阈值效应。我们提出,突变实由训练结果概率分布的连续变化驱动,当性能在随机种子间呈双峰分布时尤为明显。我们在合成长度泛化任务、多选问答和语法泛化任务中发现,不同随机种子可产生平滑或突现的缩放趋势。我们揭示,指标上的急剧突破源于种子间分布的连续变化。尽管分布可能在容量阈值处突然变为双峰,但该阈值出现在多数种子实现突破之前。即使在连续损失度量下,这一现象依然成立,说明预测模型性能时必须考虑随机性影响。

原文摘要 · Abstract (English)

Language models famously improve under a smooth scaling law, but some specific capabilities exhibit sudden breakthroughs in performance. Advocates of "emergence" view these capabilities as unlocked at a specific scale, but others attribute breakthroughs to superficial metric thresholding effects. We propose that breakthroughs are instead driven by continuous changes in the probability distribution of training outcomes when performance is bimodally distributed across random seeds. we show that different random seeds can produce either smooth or emergent scaling trends in synthetic length generalization tasks, multiple choice question answering, and grammatical generalization. We reveal that sharp breakthroughs in metrics are produced by underlying continuous changes in their distribution across seeds. These distributions may become abruptly bimodal at a capacity threshold but this threshold appears at scales well before most seeds achieve breakthrough. Our observations hold true even under continuous loss metrics, confirming that random variation must be considered when predicting a model's performance from its scale.

大模型能力突现随机性可复现性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。