arXiv:2507.19247cs.LGcs.AI2025-07被引 2

用马尔可夫范畴理论解释语言模型生成机制与训练原理。

A Markov Categorical Framework for Language Modeling

  • 将生成过程拆解为信息处理阶段,构建可组合的分析框架。
  • 揭示负对数似然训练能学习到条件不确定性与预测原型。
  • 适用于理解模型内部几何结构与高效推理方法设计。

自回归语言模型表现卓越,但其内部机制、训练如何塑造表示以及这些表示为何支持复杂行为,尚缺乏统一理论。本文提出一个基于马尔可夫范畴的分析框架,将单步生成建模为信息处理阶段的组合。该框架连接了三个常被孤立研究的方面:训练目标、学习表征空间的几何结构、以及实际模型能力。首先,框架为并行草稿方法(如推测解码)提供了信息论依据,量化隐藏状态中关于未来词元的超额信息量(超出下一个词元)。其次,阐明标准负对数似然(NLL)目标不仅学习最可能的下一个词元,还学习数据固有的条件不确定性,以范畴熵形式表达。主要谱定理为条件性结果:对于具有有界输出特征的线性-软最大头,经过白化或方差归一化后,校准的二次上界代理损失会诱导出广义的CCA/特征值问题,使表征方向与预测原型对齐。这为理解信息在模型中的流动方式及似然训练如何塑造内部几何提供了组合视角。

原文摘要 · Abstract (English)

Autoregressive language models achieve remarkable performance, yet a unified theory explaining their internal mechanisms, how training shapes representations, and why these representations support complex behavior remains incomplete. We introduce an analytical framework that models the single-step generation process as a composition of information-processing stages using the language of Markov categories. This compositional perspective connects three aspects of language modeling that are often studied separately: the training objective, the geometry of the learned representation space, and practical model capabilities. First, our framework gives an information-theoretic rationale for parallel drafting methods such as speculative decoding by quantifying the information surplus a hidden state contains about future tokens beyond the immediate next one. Second, we clarify how the standard negative log-likelihood (NLL) objective learns not only a most likely next token, but also the data's intrinsic conditional uncertainty, formalized through categorical entropy. Our main spectral result is conditional: for a linear-softmax head with bounded output features, a calibrated quadratic upper-bound surrogate to NLL induces, after whitening or variance normalization, a generalized CCA/eigenproblem aligning representation directions with predictive prototypes. This gives a compositional lens for understanding how information flows through a model and how likelihood training can shape its internal geometry.

语言模型信息论表征几何

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。