统一解释序列模型的系数动态,揭示其设计规律与性能权衡。
Design Principles for Sequence Models via Coefficient Dynamics
- 将输出系数建模为受脉冲输入驱动的线性动力系统。
- 揭示注意力机制与RNN/SSM在数学上的共性结构。
- 提供可指导新架构设计的表达能力与稳定性准则。
深度序列模型(如Transformer、状态空间模型和门控线性RNN)本质上将输出计算为过去值向量的线性组合。为深入理解并系统比较这些架构,我们提出一个统一框架,明确展现这一输出操作:将线性组合系数视为由脉冲输入驱动的自主线性动力系统的输出。该视角与以往关注线性RNN与线性注意力关联的研究显著不同,揭示了多种架构间的共同数学主题,并关键地涵盖了软最大注意力机制,适用于RNN、SSM及相关模型。不同于通常在基准上评估的新模型,我们推导出连接架构选择与模型性质的设计原则,识别出表达能力与高效实现之间的权衡、输入选择性的几何约束,以及数值稳定训练和信息保留的稳定性条件。通过整合近期文献中的多个见解,该框架既解释了现有设计的实证成功,也为系统化设计新型序列模型架构提供了指导原则。
原文摘要 · Abstract (English)
Deep sequence models, ranging from Transformers and State Space Models (SSMs) to more recent approaches such as gated linear RNNs, fundamentally compute outputs as linear combinations of past value vectors. To draw insights and systematically compare such architectures, we develop a unified framework that makes this output operation explicit, by casting the linear combination coefficients as the outputs of autonomous linear dynamical systems driven by impulse inputs. This viewpoint, in spirit substantially different from approaches focusing on connecting linear RNNs with linear attention, reveals a common mathematical theme across diverse architectures and crucially captures softmax attention, on top of RNNs, SSMs, and related models. In contrast to new model proposals that are commonly evaluated on benchmarks, we derive design principles linking architectural choices to model properties. Thereby identifying tradeoffs between expressivity and efficient implementation, geometric constraints on input selectivity, and stability conditions for numerically stable training and information retention. By connecting several insights and observations from recent literature, the framework both explains empirical successes of recent designs and provides guiding principles for systematically designing new sequence model architectures.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。