arXiv:2410.04279cs.LGstat.ML2024-10被引 1

深度网络训练可转化为凸优化问题,揭示其对称性本质

Black Boxes and Looking Glasses: Multilevel Symmetries, Reflection Planes, and Convex Optimization in Deep Networks

  • 用几何代数将绝对值激活的深度网络转为等价凸Lasso问题
  • 深度网络天然偏好对称结构,深度越大对称层级越多
  • 特征对应反射超平面距离,适合研究模型内部几何结构

我们证明,使用绝对值激活函数且输入维度任意的深度神经网络(DNN)训练问题,可等价转化为基于几何代数表达的新特征的凸Lasso问题。该形式揭示了网络中编码的对称性几何结构。通过等价的Lasso形式,我们严格证明了深度与浅层网络的根本差异:深度网络在拟合函数时天然倾向于对称结构,且深度增加可实现多层级对称(即对称中的对称)。此外,Lasso特征表示到反射超平面的距离,这些超平面由训练数据张成,并垂直于最优权重向量。数值实验验证了理论预测,展示了在使用大语言模型生成嵌入训练网络时的理论特征。

原文摘要 · Abstract (English)

We show that training deep neural networks (DNNs) with absolute value activation and arbitrary input dimension can be formulated as equivalent convex Lasso problems with novel features expressed using geometric algebra. This formulation reveals geometric structures encoding symmetry in neural networks. Using the equivalent Lasso form of DNNs, we formally prove a fundamental distinction between deep and shallow networks: deep networks inherently favor symmetric structures in their fitted functions, with greater depth enabling multilevel symmetries, i.e., symmetries within symmetries. Moreover, Lasso features represent distances to hyperplanes that are reflected across training points. These reflection hyperplanes are spanned by training data and are orthogonal to optimal weight vectors. Numerical experiments support theory and demonstrate theoretically predicted features when training networks using embeddings generated by Large Language Models.

深度学习凸优化对称性几何代数

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。