arXiv:2502.05300cs.LGcond-mat.dis-nn2025-02被引 13

用参数对称性统一解释深度学习中的层级学习现象

Parameter Symmetry Potentially Unifies Deep Learning Theory

  • 提出参数对称性破缺与恢复是统一学习动态的核心机制
  • 连接学习动力学、模型复杂度和表征形成三大层次
  • 为现代AI提供类似物理理论的统一原理,适合理论研究者

现代大型AI系统的学习过程具有层级性,常表现出类似物理系统相变的突变、定性跃迁。尽管这些现象有望揭示神经网络与语言模型的内在机制,现有理论仍零散,仅针对特定情况。本文主张,参数对称性研究方向在统一这些理论中至关重要。核心假设是:参数对称性破缺与恢复,是人工智能模型层级学习行为的统一机制。我们整合已有观察与理论,论证该方向或可建立对神经网络三个不同层级——学习动力学、模型复杂度、表征形成——的统一理解。通过连接这些层级,本文将对称性——理论物理的基石——提升为现代人工智能潜在的根本原则。

原文摘要 · Abstract (English)

The dynamics of learning in modern large AI systems is hierarchical, often characterized by abrupt, qualitative shifts akin to phase transitions observed in physical systems. While these phenomena hold promise for uncovering the mechanisms behind neural networks and language models, existing theories remain fragmented, addressing specific cases. In this position paper, we advocate for the crucial role of the research direction of parameter symmetries in unifying these fragmented theories. This position is founded on a centralizing hypothesis for this direction: parameter symmetry breaking and restoration are the unifying mechanisms underlying the hierarchical learning behavior of AI models. We synthesize prior observations and theories to argue that this direction of research could lead to a unified understanding of three distinct hierarchies in neural networks: learning dynamics, model complexity, and representation formation. By connecting these hierarchies, our position paper elevates symmetry -- a cornerstone of theoretical physics -- to become a potential fundamental principle in modern AI.

深度学习理论对称性学习动力学统一框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。