揭示神经网络训练中概念线性表示的动态形成机制
How are linear representations learned? Exact solutions to the dynamics of abstraction

- 构建线性网络精确解,解析概念方向随训练演化规律
- 发现网络深度越大、初始化尺度越小,抽象程度越高
- 理论可推广至非线性网络,指导大模型可解释性提升
在人工与生物神经网络中,概念常以表征空间中的稳定线性方向编码。深度学习中的线性表征假说支撑了基于线性探测器的概念检测与激活操控等可解释性方法。然而,以往研究多关注训练后是否存在此类方向,而对训练过程中方向如何形成的动态机制仍不清晰。本文提出“抽象”这一概念,系统研究训练中概念方向的对齐过程。在最小线性网络设定下,我们获得抽象轨迹的精确解,揭示三大核心规律:(i) 数据与目标几何共同决定最终抽象程度;(ii) 抽象随网络深度增加而增强;(iii) 初始化尺度控制训练中可达的最大抽象水平。将理论拓展至非线性网络,分析不同非线性函数的影响:erf网络近似线性理论,而ReLU网络的抽象更依赖输入几何而非目标几何。我们进一步证明一个关键衰减定律:两类非线性均使激活层的抽象程度弱于预激活层。该定律在DINOv3和Gemma 4等开放模型中得到验证,并成功用于提升大语言模型中线性探测器的泛化能力。本研究建立了一套抽象的动力学理论,对可解释性与可控性具有重要启示。
原文摘要 · Abstract (English)
In artificial and biological neural networks, concepts are often encoded as consistent linear directions in representation space. In deep learning, this idea is known as the linear representation hypothesis and underpins many interpretability and control methods based on linear probes, from concept detection to activation steering. Yet while prior work has studied whether such directions should exist $\textit{after}$ training, the dynamics of how they emerge $\textit{during}$ training remain poorly understood. Here, we develop a framework to study the alignment of concept directions during training - a process we call "abstraction". In a minimal linear network setting, we obtain exact solutions for the full trajectory of abstraction. These solutions reveal key analytic principles governing abstraction: (i) data and target geometry jointly determine abstraction at the end-of-learning, (ii) abstraction improves with network depth, and (iii) initialization scale controls the maximum abstraction reached during training. Extending our theory to nonlinear networks, we analyze how the choice of nonlinearity affects abstraction dynamics: erf networks approximate the linear theory, while abstraction in ReLU networks depends less on target geometry and more on input geometry. Across both, we prove a striking attenuation law: both nonlinearities weaken abstraction in activations relative to preactivations. We find evidence for this law in open models (DINOv3, Gemma 4) and apply our theory to improve linear probe generalization in LLMs. Together, our results provide a dynamical theory of abstraction with implications for interpretability and control.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。