arXiv:2411.03611stat.MLcs.LG2024-11

用广义熵构造线性化势函数,实现神经网络优化的指数收敛

Designing a Linearized Potential Function in Neural Network Optimization Using Csiszár Type of Tsallis Entropy

  • 基于蔡斯查尔型塔利斯熵设计线性化势函数
  • 首次在广义熵框架下实现指数收敛性证明
  • 适用于追求快速收敛的深度学习优化研究者

近年来,神经网络的学习可视为概率测度空间中的优化问题。为实现对最优解的指数收敛,基于香农熵的正则项起关键作用。尽管熵函数对收敛性影响重大,其推广却因两大技术难题而受限:一是广义对数索博列夫不等式的充分条件不足,二是梯度流方程中势函数依赖分布。本文提出利用蔡斯查尔型塔利斯熵构建线性化势函数的新框架,并证明该框架能导出指数收敛结果。

原文摘要 · Abstract (English)

In recent years, learning for neural networks can be viewed as optimization in the space of probability measures. To obtain the exponential convergence to the optimizer, the regularizing term based on Shannon entropy plays an important role. Even though an entropy function heavily affects convergence results, there is almost no result on its generalization, because of the following two technical difficulties: one is the lack of sufficient condition for generalized logarithmic Sobolev inequality, and the other is the distributional dependence of the potential function within the gradient flow equation. In this paper, we establish a framework that utilizes a linearized potential function via Csiszár type of Tsallis entropy, which is one of the generalized entropies. We also show that our new framework enable us to derive an exponential convergence result.

优化理论熵正则化指数收敛

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。