arXiv:2507.05035cs.LG2025-07被引 1

用NTK分析网络内部动态,发现性能提升不等于机制理解

Beyond Scaling Curves: Internal Dynamics of Neural Networks Through the NTK Lens

  • 通过NTK视角研究模型与数据缩放下的内部行为
  • 相同性能增长指数下,内部动态可能完全相反
  • 揭示有限宽度下特征学习的极限,适合研究模型机制者

缩放定律为神经网络性能与计算成本的关系提供了重要见解,但其内在机制仍不清楚。本文通过神经正切核(NTK)视角,实证分析了在数据和模型缩放下神经网络的行为。研究发现,尽管标准视觉任务中性能缩放指数相似,内部模型动态却可能截然相反,表明仅靠性能缩放无法揭示网络底层机制。同时,我们解决了神经网络缩放中一个未解问题:收敛到无限宽极限如何影响有限宽度模型的缩放行为。通过研究模型宽度增加时特征学习的衰减,量化了核驱动与特征驱动缩放阶段的过渡。在实验设置中,支持特征学习的最大模型宽度超过十倍于典型大语言模型的宽度。

原文摘要 · Abstract (English)

Scaling laws offer valuable insights into the relationship between neural network performance and computational cost, yet their underlying mechanisms remain poorly understood. In this work, we empirically analyze how neural networks behave under data and model scaling through the lens of the neural tangent kernel (NTK). This analysis establishes a link between performance scaling and the internal dynamics of neural networks. Our findings of standard vision tasks show that similar performance scaling exponents can occur even though the internal model dynamics show opposite behavior. This demonstrates that performance scaling alone is insufficient for understanding the underlying mechanisms of neural networks. We also address a previously unresolved issue in neural scaling: how convergence to the infinite-width limit affects scaling behavior in finite-width models. To this end, we investigate how feature learning is lost as the model width increases and quantify the transition between kernel-driven and feature-driven scaling regimes. We identify the maximum model width that supports feature learning, which, in our setups, we find to be more than ten times smaller than typical large language model widths.

神经网络缩放定律NTK

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。