揭示输入与任务如何动态塑造神经网络内部表示的幂律各向异性
Learning reshapes power-law anisotropy in internal representations

- 通过解析求解两层线性网络学习过程,发现特征学习阶段幂律指数随训练非单调变化
- 在特征学习区,内部表示谱的局部幂律指数可呈现多达四种渐近行为模式
- 结果在非线性网络中也成立,适用于理解真实模型中复杂表征的形成机制
高维信息处理系统中内部表示的幂律各向异性在从先进语言模型到小鼠大脑皮层的多种生物与人工神经网络中普遍存在。这一几何特性是理论分析的基础,但其如何由输入结构和任务驱动的学习过程产生仍不明确。本文通过精确求解具有幂律输入与教师结构的宽两层线性神经网络在教师-学生框架下的学习动力学,发现:在特征学习阶段,内部表示谱的局部幂律指数随训练过程非单调演变,并在不同模式与训练时间下呈现最多四个不同的渐近区域;而在懒惰学习区,指数基本保持不变。进一步数值实验表明,类似指数动态同样出现在更现实的非线性网络中。这些结果共同揭示了一种普遍机制——输入统计特性与任务结构的动态相互作用,催生了幂律内部表示。
原文摘要 · Abstract (English)
Power-law anisotropy in internal representations has been observed across a wide range of biological and artificial neural systems, from state-of-the-art language models to the mouse cerebral cortex. This anisotropy is a key geometric property of high-dimensional information processing and underlies a variety of theoretical analyses. However, the mechanism by which it emerges from input structure and task-driven learning has remained unclear. Here, we characterize this formation process by exactly solving the learning dynamics of a wide two-layer linear neural network in a teacher--student setting with power-law input and teacher structures. We show that, in the feature-learning regime, the local power-law exponent of the internal-representation spectrum evolves nonmonotonically over the course of training and exhibits up to four distinct asymptotic regimes across modes and training times. By contrast, in the lazy regime, the exponent remains essentially unchanged. We further demonstrate numerically that similar exponent dynamics arise in more realistic nonlinear networks. Together, these results suggest a general mechanism by which the dynamic interaction between input statistics and task structure gives rise to power-law internal representations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。