arXiv:2510.01878cs.LG2025-10被引 1

用几何原理优化大模型训练,让随机降维更高效可靠

Geometrically Principled Randomized Optimization for Efficient LLM Training

  • 基于梯度子空间几何特性设计随机降维方法
  • 在多个大模型上实现当前最佳训练效果
  • 适合关注高效训练与理论创新的研究者

大语言模型的低秩梯度优化分为结构化方法和随机方法。本文质疑随机投影为何有效,发现其根源在于梯度子空间具有近似平坦的优化景观,且大部分梯度信息位于核心子空间之外。基于此,结合随机线性代数理论,我们证明随机低秩投影可保留几何结构,并提出GrassWalk与GrassJump算法,在格拉斯曼流形上通过随机游走与跳跃进行探索。结合子空间感知优化器与梯度信号恢复,我们在LLaMA-1B、LLaMA-7B和Qwen-1.5B预训练任务中取得当前最优结果。研究重新定义随机化不仅是计算捷径,更是高维优化中的几何合理方法。

原文摘要 · Abstract (English)

Low-rank gradient optimization for large language models is currently divided into two categories: structured methods that rigorously identify subspaces, and randomized approaches employed primarily for computational efficiency. In this work, we question the intuition behind why random projections are effective. We trace this phenomenon to the geometry of the gradient subspaces, which exhibits subspace optimization landscape has a nearly flat curvature, while a significant portion of gradient information lies outside the core subspace. Leveraging these insights, and drawing on randomized linear algebra, we theoretically establish that random low-rank projections preserve the geometry, and we introduce GrassWalk and GrassJump, algorithms that navigate the Grassmannian manifold via random walks and jumps. By coupling this randomized exploration with subspace-aware optimizer and recovering the lost gradient signals, we achieve state-of-the-art results on LLaMA-1B, LLaMA-7B, and Qwen-1.5B pretraining. Our findings reframe randomization not merely as a computational shortcut, but as a geometrically principled approach to high-dimensional optimizations.

大模型训练随机优化几何学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。