arXiv:2503.08748cs.LGcs.AI2025-03被引 4

用广义熵构造新优化算法,提升收敛性与稳定性。

Mirror Descent and Novel Exponentiated Gradient Algorithms Using Trace-Form Entropies and Deformed Logarithms

  • 基于变形对数定义的广义熵设计新优化方法
  • 在非欧几何下保持统计结构,收敛更快更稳
  • 适合需要自适应优化的复杂模型训练

本文提出一类基于变形对数定义的迹形式熵的镜面下降(MD)与广义指数梯度(GEG)算法。利用这些广义熵,所提出的算法在收敛性、对梯度消失/爆炸的鲁棒性以及通过镜映映射适应非欧几里得几何方面表现更优。我们揭示了这些方法与Amari自然梯度之间的深层联系,建立了一种统一的几何基础,涵盖加法、乘法与自然梯度更新。聚焦于Tsallis、Kaniadakis、Sharma--Taneja--Mittal和Kaniadakis--Lissia--Scarfone熵族,我们证明每种熵在参数空间上诱导出不同的黎曼度量,使对应的GEG算法保持自然统计几何结构。变形对数的可调参数支持自适应几何选择,在经典欧氏优化基础上实现更强的鲁棒性与收敛性。整体框架将关键的一阶镜面下降优化方法统一于基于广义Bregman散度的信息几何视角下,其中熵的选择决定底层度量与对偶几何结构。

原文摘要 · Abstract (English)

This paper introduces a broad class of Mirror Descent (MD) and Generalized Exponentiated Gradient (GEG) algorithms derived from trace-form entropies defined via deformed logarithms. Leveraging these generalized entropies yields MD \& GEG algorithms with improved convergence behavior, robustness to vanishing and exploding gradients, and inherent adaptability to non-Euclidean geometries through mirror maps. We establish deep connections between these methods and Amari's natural gradient, revealing a unified geometric foundation for additive, multiplicative, and natural gradient updates. Focusing on the Tsallis, Kaniadakis, Sharma--Taneja--Mittal, and Kaniadakis--Lissia--Scarfone entropy families, we show that each entropy induces a distinct Riemannian metric on the parameter space, leading to GEG algorithms that preserve the natural statistical geometry. The tunable parameters of deformed logarithms enable adaptive geometric selection, providing enhanced robustness and convergence over classical Euclidean optimization. Overall, our framework unifies key first-order MD optimization methods under a single information-geometric perspective based on generalized Bregman divergences, where the choice of entropy determines the underlying metric and dual geometric structure.

优化算法信息几何镜面下降广义熵

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。