发现神经网络学习中自发出现对称性,可大幅压缩模型并提升持续学习能力。
Emergence of Fibrations, Compression, and Symmetry Breaking in Artificial Neural Networks

- 通过图论中的覆盖对称性解释深层网络学习机制
- 模型压缩至原大小17%仍保持性能,且在持续学习中表现更优
- 为黑箱模型提供可解释性,适合关注模型效率与可解释性的研究者
人工神经网络常被视为强大但不可解释的黑箱。本文证明,深度网络的学习过程会生成图论中称为纤维化和覆盖对称性的局部对称性。我们证明覆盖对称性是随机梯度下降的稳定吸引子。实验显示,这一对称性在多层、卷积、循环及Transformer等主流网络架构中普遍存在。利用这些对称性可实现显著模型压缩——将网络缩减至原始大小的17%而不损失性能。此外,受控打破覆盖对称性可缓解可塑性下降问题,在持续学习任务中达到当前最优表现。理论成果为基于对称性的智能系统提供了新基础,使黑箱变为可解释的彩色图,并支持更高效的推理与终身学习。
原文摘要 · Abstract (English)
Artificial neural networks are often regarded as powerful yet opaque black boxes. Here, we demonstrate that learning in deep neural networks generates local symmetries known in graph theory as fibrations and coverings. We prove that covering symmetries are stable attractors of stochastic gradient descent. Consistent with this theory, we report the emergence of covering symmetries across major network architectures, including multilayer, convolutional, recurrent, and transformer networks. Exploiting these symmetries enables drastic model compression - reducing networks to 17% of their original size without sacrificing performance. Furthermore, controlled breaking of covering symmetry overcomes the loss of plasticity, achieving state-of-the-art performance in continual learning. The theoretical results provide a new foundation for AI systems based on symmetries that convert black boxes into interpretable colored graphs and enable more efficient inference and lifelong learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。