arXiv:2603.08651cs.LGhep-th2026-03

用群熵理论扩展镜面下降法,实现更灵活的优化更新。

Group Entropies and Mirror Duality: A Class of Flexible Mirror Descent Updates for Machine Learning

  • 基于群熵构造可调的镜面映射,通过多参数对数函数实现自适应更新。
  • 在单纯形约束的二次规划中验证,算法具优异鲁棒性与收敛性。
  • 适合需要定制优化器的机器学习场景,如分布不均衡数据建模。

本文提出一个融合群论与群熵的理论与算法框架,构建无限且灵活的镜面下降(MD)优化算法族。该方法利用群熵这一由群合成法则定义的广义熵泛函,涵盖并显著拓展了所有迹形式熵(如香农、塔斯利斯、卡尼亚达基斯族)。通过群论镜面映射(即链接函数),以多参数广义对数及其逆(群指数)表达,实现对数据几何结构和统计分布的高度适应性。引入“镜面对偶”概念,在特定学习率约束下可无缝切换链接函数与其逆,实现灵活更新。通过调节群对数超参数,可适配训练分布的统计特性,并同时保证良好收敛性。该通用性不仅提升灵活性与收敛性能,还为正则化设计与自然梯度算法提供新思路。在大规模单纯形约束二次规划问题上进行了广泛评估,验证了所提更新的有效性、鲁棒性与性能优势。

原文摘要 · Abstract (English)

We introduce a comprehensive theoretical and algorithmic framework that bridges formal group theory and group entropies with modern machine learning, paving the way for an infinite, flexible family of Mirror Descent (MD) optimization algorithms. Our approach exploits the rich structure of group entropies, which are generalized entropic functionals governed by group composition laws, encompassing and significantly extending all trace-form entropies such as the Shannon, Tsallis, and Kaniadakis families. By leveraging group-theoretical mirror maps (or link functions) in MD, expressed via multi-parametric generalized logarithms and their inverses (group exponentials), we achieve highly flexible and adaptable MD updates that can be tailored to diverse data geometries and statistical distributions. To this end, we introduce the notion of \textit{mirror duality}, which allows us to seamlessly switch or interchange group-theoretical link functions with their inverses, subject to specific learning rate constraints. By tuning or learning the hyperparameters of the group logarithms enables us to adapt the model to the statistical properties of the training distribution, while simultaneously ensuring desirable convergence characteristics via fine-tuning. This generality not only provides greater flexibility and improved convergence properties, but also opens new perspectives for applications in machine learning and deep learning by expanding the design of regularizers and natural gradient algorithms. We extensively evaluate the validity, robustness, and performance of the proposed updates on large-scale, simplex-constrained quadratic programming problems.

优化算法镜面下降群熵自适应更新

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。