arXiv:2607.22757cs.LGcs.AI2026-07

给大模型引入分层评分机制,提升表达能力且不影响推理效率。

Hierarchical Grading in Large Language Models

  • 用代数框架在嵌入空间中引入分层结构,通过加权标量作用传播至注意力和训练目标。
  • 最优分层配置为凸锥内的闭式解,比普通Transformer更优且具理论保障。
  • 可离线计算最优分层参数,训练后仍保持标准Transformer架构与速度。

我们提出分级大语言模型(GLLMs),一种代数框架,为Transformer的表示空间引入分层结构,并将诱导的加权标量作用传播至嵌入、自注意力及训练目标。该构造扩展了分级神经网络与分级Transformer的理论至自回归语言模型,同时保持表达能力、渐近计算复杂度与推理开销。其几何基础为几何不变理论:分层的优势由分级环上的Kempf-Ness泛函刻画;改善均匀架构的分层构成一个开凸锥,其成员由希尔伯特-穆尔福德型判据决定,即方向与目标和数据的两个可测轮廓配对;最优分层是两个动量映射的重合点,具有闭式表达;普通Transformer仅为该凸锥边界上一个半稳定各向同性点,属于更大的分级家族而非最优解。对于层级分层目标,我们证明了分级先验与无分级先验之间的极小极大分离:在所有估计器下,两类目标的风险在样本量的显式窗口内分离,且随层级数呈指数衰减。两个轮廓均可离线估计,因此最优分层可通过预训练前的凸规划求解。由于分层在训练后被吸收进参数,每个GLLM均可编译为相同架构与推理复杂度的标准Transformer。

原文摘要 · Abstract (English)

We introduce Graded Large Language Models (GLLMs), an algebraic framework that equips the representation space of a transformer with a grading and propagates the induced weighted scalar action through embeddings, self-attention, and the training objective. The construction extends the theory of graded neural networks and graded transformers to autoregressive language models while preserving expressive power, asymptotic computational complexity, and inference cost. The governing geometric picture is that of geometric invariant theory. The benefit of a grading is expressed by a Kempf--Ness functional on the grading torus; the grades that improve upon the uniform architecture form an open convex cone whose membership is decided by a Hilbert--Mumford-type criterion pairing a grade direction against two measurable profiles of the target and the data; the optimal grades are the coincidence point of two moment maps, given in closed form; and the ordinary transformer appears as a semistable isotropic point on the boundary of the cone: one member of a larger graded family rather than a distinguished optimum. Separately, for level-stratified targets we prove a minimax separation between the graded prior and its absence: over all estimators the risks of the graded and uniform target classes separate throughout an explicit window of sample sizes, by a factor that decays exponentially in the number of levels under geometric stratification. Both profiles are estimable offline, so the optimal grades solve a convex program certified before training begins. Because the grading is absorbed into the learned parameters after training, every GLLM compiles to a standard transformer of identical architecture and inference complexity.

大模型分层结构代数框架优化理论

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。