提出可实时检测神经网络概念涌现的相变理论,揭示特征形成的数学机制。
Feature Lottery? A Bifurcation Theory of Concept Emergence

- 基于损失海森矩阵构建动态相坐标β/β_c,实现无标签实时监测
- 发现特征结构出现对应超临界劈裂分岔,早期5%训练即可预测最终纯度
- 解释了模型延迟逃逸现象,为训练监控与调试提供早期预警信号
神经网络在训练过程中会于特定时刻形成结构化表征,但传统检测方法依赖事后、有标签指标。本文提出表征动力学的分岔理论,通过附加于编码器的被动高斯混合模型探测器,发现结构出现对应由损失海森矩阵驱动的超临界劈裂分岔。系统存在理论上可预测的零交叉点(β_c),其与当前状态(β)的比值β(t)/β_c(t)构成全由隐藏状态计算的通用无标签相坐标。实证验证了该坐标在语言模型SAE(Pythia)、自监督学习(CIFAR)及grokking(模运算)中预测的四种过渡态。关键发现:有限耗散下宏观对称性破缺可滞后零交叉数个数量级,为grokking延迟逃逸提供了严格动力学解释。微观上,分岔产生共享不稳定子空间,强制集体对称性破缺。我们称之为SAE训练中的「特征彩票」:特征的最终可解释性可在极早期预测——训练仅5%时,早期原子纯度即能稳健预测收敛纯度,前十分位早期原子在收敛时纯度达基线12倍以上。该坐标不仅解释概念涌现,还可作为训练健康度的实用早期预警,提前检测可用结构出现、特征身份固化及表征崩溃阶段。
原文摘要 · Abstract (English)
Neural networks acquire structured representations at specific moments during training, yet identifying these transitions typically relies on retrospective, label-dependent metrics. We introduce a bifurcation theory of representation dynamics to detect these moments in real time. Analyzing a passive GMM probe attached to the evolving encoder, we show the onset of structure corresponds to a supercritical pitchfork bifurcation driven by the loss Hessian. The system exhibits a theoretically predictable zero-crossing ($β_c$) that, compared to the network's current state ($β$), yields a dynamic ratio $β(t)/β_c(t)$: a universal, label-free phase coordinate for representation dynamics, computable entirely from hidden states. We empirically validate four distinct transition regimes predicted by this coordinate across diverse settings: SAEs on language models (Pythia), SSL (CIFAR), and grokking (modular arithmetic). Crucially, under finite dissipation, macroscopic symmetry-breaking can lag the initial zero-crossing by orders of magnitude, which providing a rigorous dynamical account of the delayed escape observed in grokking. Microscopically, the bifurcation creates a shared unstable subspace, forcing collective symmetry breaking. We term this the "feature lottery" in SAE training: a feature's terminal interpretability becomes predictable remarkably early. By only 5% of training, early atom purity robustly predicts final convergence purity, with top-decile early atoms achieving over 12x the baseline purity at convergence. Beyond explaining concept emergence, $β/β_c$ provides a practical early-warning indicator for training health, detecting the onset of usable structure, the crystallization of feature identity, and representational collapse epochs before downstream metrics react.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。