提出统一优化框架,让机器学习在非欧几何下更高效
A Universal Banach--Bregman Framework for Stochastic Iterations: Unifying Stochastic Mirror Descent, Learning and LLM Training
- 用巴拿赫-布雷格曼几何统一镜像下降、自然梯度等方法
- 在非希尔伯特空间实现超松弛加速,收敛速度提升20%
- 适用于大模型训练,对复杂任务有更强鲁棒性
随机优化支撑现代人工智能的可扩展性,涵盖机器学习、深度学习、强化学习和大语言模型训练。然而现有理论大多局限于希尔伯特空间,依赖内积框架与正交性,无法捕捉非欧几里得场景,如单纯形上的镜像下降、稀疏学习中的Bregman近似、信息几何中的自然梯度,或基于KL散度正则的大语言模型训练。本文首次引入巴拿赫-布雷格曼框架,以一般巴拿赫空间为基础,建立统一模板:通过Bregman投影与Bregman-Féjer单调性涵盖随机逼近、镜像下降、自然梯度、自适应方法及镜像近端;在非希尔伯特空间中实现超松弛(λ > 2),揭示其加速机制;并提供从几乎必然有界到几何收敛率的完整收敛定理。在合成数据和真实任务上验证有效,包括UCI基准、Transformer训练、演员-评论家强化学习,以及WikiText-2上使用distilGPT-2的大语言模型训练,结果表明收敛速度最快提升20%,方差降低,精度更高。该框架为下一代优化理论与实践提供了统一基础。
原文摘要 · Abstract (English)
Stochastic optimization powers the scalability of modern artificial intelligence, spanning machine learning, deep learning, reinforcement learning, and large language model training. Yet, existing theory remains largely confined to Hilbert spaces, relying on inner-product frameworks and orthogonality. This paradigm fails to capture non-Euclidean settings, such as mirror descent on simplices, Bregman proximal methods for sparse learning, natural gradient descent in information geometry, or Kullback--Leibler-regularized language model training. Unlike Euclidean-based Hilbert-space methods, this approach embraces general Banach spaces. This work introduces a pioneering Banach--Bregman framework for stochastic iterations, establishing Bregman geometry as a foundation for next-generation optimization. It (i) provides a unified template via Bregman projections and Bregman--Fejer monotonicity, encompassing stochastic approximation, mirror descent, natural gradient, adaptive methods, and mirror-prox; (ii) establishes super-relaxations ($λ> 2$) in non-Hilbert settings, enabling flexible geometries and elucidating their acceleration effect; and (iii) delivers convergence theorems spanning almost-sure boundedness to geometric rates, validated on synthetic and real-world tasks. Empirical studies across machine learning (UCI benchmarks), deep learning (e.g., Transformer training), reinforcement learning (actor--critic), and large language models (WikiText-2 with distilGPT-2) show up to 20% faster convergence, reduced variance, and enhanced accuracy over classical baselines. These results position Banach--Bregman geometry as a cornerstone unifying optimization theory and practice across core AI paradigms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。