arXiv:2603.05573cs.LG2026-03

从李代数视角揭示序列模型深度对表达能力的影响机制。

Why Depth Matters in Parallelizable Sequence Models: A Lie Algebraic View

  • 用李代数框架建立深度与模型表达力的对应关系。
  • 证明误差随深度增加呈指数下降,与实际性能一致。
  • 适合研究模型表达力、深度与并行计算权衡的学者。

可扩展的序列模型(如Transformer变体和结构化状态空间模型)常为实现序列级并行而牺牲表达能力,从而支持高效训练。本文从李代数控制视角分析模型在超出其表达力范围时的误差边界及其缩放规律。理论表明,序列模型的深度与李代数扩张层级存在一一对应关系。呼应近期理论研究,我们刻画了常数深度序列模型的李代数类别及其表达力上限。进一步,我们解析推导出近似误差界,证明误差随深度增加呈指数衰减,与这些模型的强大经验表现一致。我们在符号词和连续状态追踪任务上验证了理论预测的有效性。

原文摘要 · Abstract (English)

Scalable sequence models, such as Transformer variants and structured state-space models, often trade expressivity power for sequence-level parallelism, which enables efficient training. Here we examine the bounds on error and how error scales when models operate outside of their expressivity regimes using a Lie-algebraic control perspective. Our theory formulates a correspondence between the depth of a sequence model and the tower of Lie algebra extensions. Echoing recent theoretical studies, we characterize the Lie-algebraic class of constant-depth sequence models and their corresponding expressivity bounds. Furthermore, we analytically derive an approximation error bound and show that error diminishes exponentially as the depth increases, consistent with the strong empirical performance of these models. We validate our theoretical predictions using experiments on symbolic word and continuous-valued state-tracking problems.

序列建模李代数深度分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。