发现MLP残差网络会按输入谱结构选择性压缩信息,像物理中的重整化群过程。
Rank Collapse, Fixed Points, and the Renormalization Group Structure of MLP Residual Networks

- 通过分析合成马尔可夫序列,发现残差流有效秩随深度单调下降。
- 短相关长度输入下秩塌陷明显,长相关长度则保留全部有用信息。
- 网络在特定层间发生显著变化,其余部分接近固定点,具重整化群特征。
深度神经网络前向传播与重整化群(RG)流之间的类比虽屡被提及,但现有研究多为定性描述:深度被视为粗粒化尺度,注意力类比于配分函数,表征被认为流向固定点。然而尚无工作定义可观测的RG序参数,也未在受控输入分布变化下验证,或做出可实证的定量预测。本文研究最简可处理架构:在具有已知谱性质的合成马尔可夫链序列上训练的纯MLP残差堆叠,用于掩码标记预测。报告三项发现:(i) 训练后残差流的有效秩随深度单调下降,符合无关自由度逐步整合;(ii) 秩塌陷具选择性:在相关长度约1的序列中显著发生,而在相关长度约7的序列中则不出现(以位置层面测量,排除均值池化伪影),网络恰好保留任务相关的自由度,契合RG相关性准则;(iii) 层间核漂移集中于一到两个特定过渡层,其余网络近似固定点,符合离散固定点平台结构。这些发现构成首个基于位置层级、量化的证据,表明MLP残差网络实施了由输入分布谱结构调控的选择性粗粒化过程。
原文摘要 · Abstract (English)
The analogy between deep neural network forward passes and renormalization group (RG) flows has been repeatedly noted in the literature, but existing treatments remain qualitative: depth is described as a coarse-graining scale, attention is likened to a partition function, and representations are said to flow toward fixed points. No existing work has defined a measurable RG order parameter, tested it under controlled variation of the input distribution, or made quantitative predictions that are empirically verified. We study the simplest architecture for which the analogy is tractable: a pure MLP residual stack trained on masked token prediction over synthetic Markov chain sequences with known spectral properties. We report three findings. (i) The effective rank of the residual stream decreases monotonically with depth after training, consistent with progressive integration of irrelevant degrees of freedom. (ii) This rank collapse is selective: it occurs for chains with short correlation length approximately 1 but is absent for chains with long correlation length approximately 7, measured at the position level to control for mean-pooling artifacts. The network preserves exactly the degrees of freedom relevant to the prediction task, the content of the RG relevance criterion. (iii) Inter-layer kernel drift is concentrated at one or two specific transitions, with the remainder of the network near a fixed point, consistent with a discrete fixed-point plateau. Together these findings constitute the first quantitative, position-level evidence that MLP residual networks implement a selective coarse-graining procedure governed by the spectral structure of the input distribution.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。