arXiv:2604.02309cs.LGcs.CL2026-04被引 2

提出高效精确的流间混合参数化方法,可显著提升模型表达能力。

go-$m$HC: Direct Parameterization of Manifold-Constrained Hyper-Connections via Generalized Orthostochastic Matrices

  • 基于广义正交随机矩阵,实现对双随机矩阵集的精确且高效参数化
  • 计算复杂度为O(d³),在相同算力下比基线方法更接近完整双随机多面体
  • 适用于需要动态层连接的模型,特别适合大规模语言模型扩展

双随机矩阵可实现残差流间的可学习混合,但对双随机矩阵集(Birkhoff多面体)进行精确且高效的参数化仍是未解难题。现有精确方法随流数d呈阶乘级增长,而克罗内克分解方法虽高效但表达能力受限。本文提出一种基于广义正交随机矩阵的新颖精确参数化方法,复杂度为O(d³),并引入单个超参数s,连续介于计算高效边界与完全表达的Birkhoff多面体之间。基于流约束超连接(mHC)框架,构建go-mHC。该方法可自然融合克罗内克分解,在相近浮点运算量下显著恢复表达能力。谱分析表明,go-mHC比克罗内克基线更完整地填充Birkhoff多面体。在合成流混合任务中,go-mHC达到理论最小损失,收敛速度最快达10倍提升。我们在一个3000万参数的GPT风格模型上验证了该方法。go-mHC在表达性、效率与精确性上的优势,为将流数d作为新模型容量维度提供了实用路径。

原文摘要 · Abstract (English)

Doubly stochastic matrices enable learned mixing across residual streams, but parameterizing the set of doubly stochastic matrices (the Birkhoff polytope) exactly and efficiently remains an open challenge. Existing exact methods scale factorially with the number of streams ($d$), while Kronecker-factorized approaches are efficient but expressivity-limited. We introduce a novel exact parameterization grounded in the theory of generalized orthostochastic matrices, which scales as $\mathcal{O}(d^3)$ and exposes a single hyperparameter $s$ which continuously interpolates between a computationally efficient boundary and the fully expressive Birkhoff polytope. Building on Manifold-Constrained Hyper-Connections ($m$HC), a framework for learned dynamic layer connectivity, we instantiate this parameterization in go-$m$HC. Our method composes naturally with Kronecker-factorized methods, substantially recovering expressivity at similar FLOP costs. Spectral analysis indicates that go-$m$HC fills the Birkhoff polytope far more completely than Kronecker-factorized baselines. On synthetic stream-mixing tasks, go-$m$HC achieves the minimum theoretical loss while converging up to $10\times$ faster. We validate our approach in a 30M parameter GPT-style language model. The expressivity, efficiency, and exactness of go-$m$HC offer a practical avenue for scaling $d$ as a new dimension of model capacity.

深度学习模型架构矩阵参数化语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。