揭示神经网络如何自动生成简洁且可迁移的表征
Formation of Representations in Neural Networks
- 提出共形表征假说,解释隐藏层表征、权重和梯度的对齐机制
- 发现打破该假说会引出幂律关系,与神经坍缩等现象相关
- 为深度学习核心现象提供统一理论框架,适合研究模型内部机制者
理解神经网络的表征有助于揭开现代AI系统的黑箱。然而,复杂、结构化且可迁移的表征如何在神经网络中形成仍是个谜。基于前期成果,我们提出共形表征假说(CRH),认为六个对齐关系普遍支配大多数隐藏层的表征形成过程。在CRH下,潜在表征(R)、权重(W)和神经元梯度(G)在训练中相互对齐,暗示神经网络自然学习紧凑表征,使神经元与权重对任务无关变换保持不变。随后我们证明,当CRH被打破时,会涌现出R、W、G间的互逆幂律关系,称为多项式对齐假说(PAH)。我们给出最小假设理论,表明梯度噪声与正则化的平衡是生成共形表征的关键。CRH与PAH共同揭示了统一神经坍缩、神经特征假设等关键深度学习现象的可能性。
原文摘要 · Abstract (English)
Understanding neural representations will help open the black box of neural networks and advance our scientific understanding of modern AI systems. However, how complex, structured, and transferable representations emerge in modern neural networks has remained a mystery. Building on previous results, we propose the Canonical Representation Hypothesis (CRH), which posits a set of six alignment relations to universally govern the formation of representations in most hidden layers of a neural network. Under the CRH, the latent representations (R), weights (W), and neuron gradients (G) become mutually aligned during training. This alignment implies that neural networks naturally learn compact representations, where neurons and weights are invariant to task-irrelevant transformations. We then show that the breaking of CRH leads to the emergence of reciprocal power-law relations between R, W, and G, which we refer to as the Polynomial Alignment Hypothesis (PAH). We present a minimal-assumption theory proving that the balance between gradient noise and regularization is crucial for the emergence of the canonical representation. The CRH and PAH lead to an exciting possibility of unifying major key deep learning phenomena, including neural collapse and the neural feature ansatz, in a single framework.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。