首次推导出完整Transformer块的精确海森矩阵,揭示各组件的曲率贡献。
Closing the Curvature Gap: Full Transformer Hessians
- 基于矩阵微积分推导出完整Transformer块的闭式海森矩阵。
- 理论证明海森矩阵各块的谱范数有界,验证了优化稳定性。
- 适合研究模型优化、架构设计或二阶方法的学者参考。
尽管Transformer模型被广泛使用,其优化景观仍不清晰。已有研究仅分析了孤立自注意力机制的曲率特性,缺乏对完整Transformer模块(包含层归一化、前馈网络和残差连接相互作用)的理论刻画。本文通过严格矩阵微积分,推导出在任意二次可微损失函数下完整Transformer块的精确闭式海森矩阵,解决了这一空白。我们处理了层归一化与逐行激活的非线性,建立了海森矩阵块的显式谱范数上界。分析表明不同组件产生不同的曲率机制,并明确了特定子层的曲率贡献。实验验证显示,该公式与自动微分结果在数值精度内完全一致,并在闭式雅可比计算中实现显著加速。
原文摘要 · Abstract (English)
The optimization landscape of Transformer models remains poorly understood despite their widespread adoption. While recent studies have derived curvature properties for isolated self-attention mechanisms, a comprehensive theoretical characterization of the full Transformer block, accounting for the interactions between Layer Normalization, Feed-Forward Networks (FFNs), and residual connections, is missing. In this work, we close this gap by deriving the exact, closed-form Hessian for the complete Transformer block under arbitrary twice-differentiable loss functions. We utilize rigorous matrix calculus to handle the non-linearities of LayerNorm and row-wise activations, establishing explicit spectral norm bounds for the resulting Hessian blocks. Our analysis reveals how different architectural components contribute distinct curvature mechanisms, identifying the specific curvature contributions of particular sub-layers. Furthermore, empirical validation against automatic differentiation confirms the exactness of the derived formulas up to numerical precision and shows substantial computational speedups for the closed-form Jacobian evaluations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。