arXiv:2605.07959cs.LGmath.FA2026-05

证明了注意力层与LoRA在随机训练下的可训练性,无需数据或模型规模假设。

Convergent Stochastic Training of Attention and Understanding LoRA

  • 建立统一框架,证明注意力层和浅层网络的可训练性。
  • 在任意温和正则化下,损失函数诱导庞加莱不等式。
  • 适用于大模型微调,对数据和模型大小无额外要求。

Transformer 架构已彻底改变机器学习,注意力层的应用日益普及。对于大型模型,低秩适应(LoRA)通过参数分解训练,实现了出色的精度-规模权衡。本文通过统一框架,严格证明了在随机优化方法下此类模型的可训练性。我们证明:对于任意温和正则化,注意力层及浅层神经网络上的经验回归损失,均诱导相应的吉布斯测度满足庞加莱不等式。由此可推导出,一种模仿 SGD 的特定随机微分方程(SDE)能最小化对应损失。这是首次在注意力机制与浅层网络上不依赖数据分布或模型规模假设,证明可训练性的成果。

原文摘要 · Abstract (English)

Transformers have revolutionized machine learning and deploying attention layers in the model is increasingly standard across a myriad of applications. Further, for large models, it is common to implement Low Rank Adaptation (LoRA), whereby a factorized parameterization of them is trained, to achieve a surprisingly beneficial accuracy-size trade-off. In this work, via a unified framework we rigorously establish trainability of such models under stochastic methods. We prove that for any mild regularization, the empirical regression loss on a attention layer and LoRA on a shallow neural net, both induce Poincaré inequality for the corresponding Gibbs' measure. Then it follows via invoking recent results that a certain SDE, which mimics the SGD, minimizes the corresponding losses. In both the cases, our first-of-its-kind results of trainability on attention and nets, do not rely on any assumptions on the data or the size of the architecture.

注意力机制LoRA随机训练理论分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。