arXiv:2502.17500cs.LGcs.AI2025-02被引 3

提出广义欧拉对数,统一多种熵与优化方法,提升模型鲁棒性与灵活性。

Generalized Euler Logarithm and its Applications in Machine Learning: Natural Gradient, Backpropagation, Generalized EG, Mirror Descent and OLPS

  • 以双参数广义对数为核,统一多类信息度量与优化算法。
  • 在深度网络中构建新损失函数,实现精确反向传播与自然梯度融合。
  • 两个参数可独立控制尾部鲁棒性与局部梯度形状,适合复杂优化场景。

本文深入研究双参数广义欧拉对数及其逆函数——变形的 (a,b)-指数函数。系统阐明了保证单调性、凹性和可逆性的参数域,推导出级数与积分表示,并揭示其与广义一、二参数对数(如Tsallis、Kaniadakis、Schwämmle-Tsallis、Kaniadakis-Scarfone、Tempesta型)的显式关联。由此确立欧拉 (a,b)-对数作为广义熵与散度度量家族的统一核心。算法层面,将该对数扩展至现代机器学习与优化:引入基于欧拉对数的广义指数梯度(GEG)与镜像下降(MD)框架,其中 (a,b)-对数作为Bregman散度中的灵活链接函数;提出欧拉基广义交叉熵(GCE)损失,推导其精确反向传播公式,并与Fisher-Rao自然梯度(NG)下降无缝集成。通过分离费舍尔信息矩阵(FIM),并发展对角近似,证明两变形参数可成功解耦尾部鲁棒性与局部梯度调节。

原文摘要 · Abstract (English)

This paper investigates in depth the fundamental properties of the two-parameter generalized Euler logarithm and its inverse, the associated deformed $(a,b)$-exponential function. We systematically clarify the parameter domains that guarantee monotonicity, concavity, and invertibility, derive series and integral representations, and provide explicit links to a broad class of one- and two-parameter deformations, including Tsallis, Kaniadakis, Schwämmle--Tsallis, Kaniadakis--Scarfone, and Tempesta-type logarithms and their inverse exponentials. In this way, the Euler $(a,b)$-logarithm is established as a unifying kernel for a wide family of generalized entropies and divergence measures. On the algorithmic side, we extend applications of the Euler logarithm to modern machine learning and optimization. We introduce generalized Exponentiated Gradient (GEG) and Mirror Descent (MD) schemes in which the Euler $(a,b)$-logarithm acts as a flexible link function in the underlying Bregman divergence. In addition, we propose an Euler-based Generalized Cross-Entropy (GCE) loss for deep neural networks, derive its exact backpropagation formulas, and detail its seamless integration with Fisher-Rao Natural Gradient (NG) descent. By isolating the Fisher Information Matrix (FIM) and developing a diagonal NG approximation, we demonstrate how the two deformation parameters successfully decouple tail robustness from local gradient shaping.

优化算法广义熵自然梯度深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。