arXiv:2507.05644cs.LGcs.AI2025-07被引 2

提出新理论解释神经网络如何学习特征,比现有猜想更可靠。

The Features at Convergence Theorem: a first-principles alternative to the Neural Feature Ansatz for how networks learn representations

论文配图:The Features at Convergence Theorem: a first-principles alternative to the Neural Feature Ansatz for how networks learn representations
图 1 · 摘自论文原文
  • 从优化条件出发推导出收敛时特征的理论规律
  • 在多种任务中与实际学习结果高度一致,涵盖特殊现象
  • 为理解模型学习机制提供可证明的理论基础,适合研究者参考

深度学习的核心挑战之一是理解神经网络如何学习表征。主流方法是神经特征假设(NFA),但其缺乏理论基础,仅凭经验验证,难以判断适用边界。本文从一阶最优性条件出发,推导出收敛特征定理(FACT),作为NFA的替代。FACT不仅在收敛时与实际学习特征匹配度更高,还能解释为何NFA在多数场景成立,并复现了模块化算术中的“领悟”现象和稀疏奇偶性学习中的相变行为等关键特征学习现象。该研究将理论优化分析与经验驱动的NFA研究统一,提供了可证明且经实证支持的收敛阶段学习机制,为理解神经网络表征学习提供了一个更坚实的理论框架。

原文摘要 · Abstract (English)

It is a central challenge in deep learning to understand how neural networks learn representations. A leading approach is the Neural Feature Ansatz (NFA) (Radhakrishnan et al. 2024), a conjectured mechanism for how feature learning occurs. Although the NFA is empirically validated, it is an educated guess and lacks a theoretical basis, and thus it is unclear when it might fail, and how to improve it. In this paper, we take a first-principles approach to understanding why this observation holds, and when it does not. We use first-order optimality conditions to derive the Features at Convergence Theorem (FACT), an alternative to the NFA that (a) obtains greater agreement with learned features at convergence, (b) explains why the NFA holds in most settings, and (c) captures essential feature learning phenomena in neural networks such as grokking behavior in modular arithmetic and phase transitions in learning sparse parities, similarly to the NFA. Thus, our results unify theoretical first-order optimality analyses of neural networks with the empirically-driven NFA literature, and provide a principled alternative that provably and empirically holds at convergence.

表征学习理论分析神经网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。