arXiv:2512.04165cs.LGstat.ML2025-12被引 5

提出简单尺度分析法,预测深度网络特征学习的出现条件。

Mitigating the Curse of Detail: Scaling Arguments for Feature Learning and Sample Complexity

  • 用尺度分析替代复杂方程,快速预判特征学习何时出现
  • 复现已有结果的标度指数,并预测三层非线性网络行为
  • 适合研究深度学习理论或网络泛化机制的学者

深度学习理论中两个核心问题是解析特征学习机制与确定丰富参数下网络的隐式偏差。现有丰富特征学习理论多以高维非线性方程形式呈现,需大量计算求解。由于深度学习问题涉及众多细节,这种分析复杂性常难以避免。本文提出一种高效启发式方法,可预测特征学习模式出现的数据规模与网络宽度。该尺度分析远比精确理论简便,且能复现多种已知结果的标度指数。此外,我们对三层非线性网络与注意力头等复杂简化架构做出新预测,拓展了深度学习第一性原理理论的应用范围。

原文摘要 · Abstract (English)

Two pressing topics in the theory of deep learning are the interpretation of feature learning (FL) mechanisms and the determination of implicit bias of networks in the rich regime. Current theories of rich FL often appear in the form of high-dimensional non-linear equations, which require computationally intensive numerical solutions. Given the many details that go into defining a deep learning problem, this analytical complexity is a significant and often unavoidable challenge. Here, we propose a powerful heuristic route for predicting the data and width scales at which various patterns of FL emerge. This form of scale analysis is considerably simpler than such exact theories and reproduces the scaling exponents of various known results. In addition, we make novel predictions on complex toy architectures, such as three-layer non-linear networks and attention heads, thus extending the scope of first-principle theories of deep learning.

深度学习理论特征学习尺度分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。