arXiv:2604.23740cs.LG2026-04

用连续动力系统解释Transformer,揭示其稳定训练的内在机制。

Transformer as an Euler Discretization of Score-based Variational Flow

  • 将Transformer视为分数路径的欧拉离散化,建立理论基础。
  • 实验显示分数流指标与任务表现高度相关,揭示深度敏感性。
  • 适合对模型原理和稳定性感兴趣的从业者或研究者。

尽管Transformer在机器学习中占据主导地位,但其架构仍主要依赖启发式设计,缺乏统一的理论基础。本文提出基于分数的变分流(SVFlow),一种用于表征学习的连续时间动力系统,其状态演化遵循条件对数似然分数的变分后验加权平均,并通过变分一致性为正则化提供严谨依据。我们证明球面SVFlow的前向欧拉离散化恰好恢复Transformer结构:多头注意力通过vMF核平滑的后验近似SVFlow向量场,MoE/前馈网络以松弛的网络方式近似该向量场,残差-归一化模块实现保持球面几何的松弛再投影。这一统一解释说明了为何注意力无需显式正则化即可稳定训练,而MoE需辅助平衡损失。在预训练语言模型上进行前缀打乱实验表明,由SVFlow导出的度量与任务性能显著相关,揭示深度依赖性,并反映注意力的内在动力学。

原文摘要 · Abstract (English)

Despite the Transformer's dominance across machine learning, its architecture remains largely heuristic and lacks a unified theoretical foundation. We introduce Score-based Variational Flow (SVFlow), a continuous-time dynamical system for representation learning in which the state evolves according to a variational posterior-weighted average of conditional log-likelihood scores, and provide a principled basis for regularization through variational consistency. We show that forward Euler discretization of spherical SVFlow exactly recovers the Transformer architecture. Multi-head attention approximates SVFlow vector field via a vMF kernel-smoothed posterior, while MoE/FFN approximates it in a relaxed network-based way, and the residual-normalization block implements a relaxed retraction that maintains spherical geometry. This unification explains why attention trains stably without explicit regularization while MoE requires auxiliary balancing losses. Experiments on pre-trained language models with prefix shuffling show that SVFlow-induced metrics correlate with task performance, reveal depth-dependent sensitivity, and reflect the intrinsic dynamics of attention.

Transformer连续动态分数流理论解释

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。