arXiv:2605.29548cs.LG2026-05被引 3

大模型能学小模型学不会的任务,因资源分配更优且干扰更少。

Why Larger Models Learn More: Effects of Capacity, Interference, and Rare-Task Retention

论文配图:Why Larger Models Learn More: Effects of Capacity, Interference, and Rare-Task Retention
图 1 · 摘自论文原文
  • 大模型通过减少任务间干扰,更高效分配神经元资源。
  • 只有大模型能学会低频高复杂度任务,小模型会忽略这些任务。
  • 适合研究模型容量、训练数据分布与任务学习的关系。

更大的模型能学习小模型无法掌握的任务。我们提出一个基于幂律缩放的简单现象学解释:即使在无限训练数据下,大模型也能学习到小模型未能捕捉的数据分布部分。通过合成任务实验发现,小模型会将神经元集中于高频或低复杂度任务,导致对稀有复杂任务表现差,即便这些任务的解存在。大模型则通过降低梯度干扰,使常见任务更新变弱,从而保留稀有任务特征的逐步积累。我们在OLMo模型(4M至4B参数)上验证了这一现象:仅大模型能学会低频复杂任务,其表征中嵌入更多任务特征,且任务间梯度干扰更小。本研究从数据驱动角度解释了为何大模型学习能力更强,为模型规模选择和数据混合策略提供依据。

原文摘要 · Abstract (English)

Larger models learn tasks smaller models do not. What drives this phenomenon? We develop a simple phenomenological argument that power-law scaling already suggests that a larger model will be able to learn a part of the data distribution that a smaller model fails to learn, even with infinite training data. To validate this claim and identify its causes, we study the effects of model scaling on a synthetic setup consisting of a mixture of tasks that show monotonic scaling curves. The results point to a data-induced competition over resources (neurons). Specifically, smaller models allocate their neurons to high frequency or low complexity tasks, and so they learn solutions that perform poorly on rare and complex tasks. Moreover, this happens even when solutions capable of expressing the desired task exist. We then assess how a larger model circumvents this data-centric bottleneck, finding that it traces to a reduced interference mechanism: larger models can allocate enough resources to common tasks that the gradient updates for those tasks become weak, which means that they do not overwrite rare-task features as they slowly accumulate. Finally, to further validate these claims, we pretrain OLMo models (4M to 4B parameters) on novel tasks of varying frequency and complexity. The results mirror those from our synthetic data experiments: only the larger OLMo models learn the infrequent and complex tasks, and these larger models embed more task features in their representations and show less gradient interference between tasks. Overall, we offer a data-centric account of why larger models learn tasks that smaller models fail to. This helps explain why larger models are better in practice, and it can inform practical questions concerning model sizing and training data mixtures.

模型容量任务学习梯度干扰

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。