解析Transformer深度与宽度对上下文学习的影响,揭示最优模型结构。
Theory of Scaling Laws for In-Context Regression: Depth, Width, Context and Time
- 构建深度线性自注意力模型研究上下文学习的缩放规律。
- 深度仅在上下文有限时提升性能,随机旋转协方差下深度效果显著。
- 给出可计算的渐近风险与幂律关系,适用于算力优化设计。
我们研究了深度线性自注意力模型中线性回归的上下文学习(ICL),分析性能如何依赖于宽度、深度、训练步数、批大小及每上下文数据量等计算与统计资源。在数据维度、上下文长度和残差流宽度成比例增长的联合极限下,考察三种ICL设置:(1) 各向同性协方差与任务(ISO),(2) 固定且结构化协方差(FS),(3) 协方差随机旋转且有结构(RRS)。在ISO与FS设置中,深度仅在上下文长度受限时有益;而在RRS设置中,即使上下文长度无限,增加深度仍能显著提升性能。该模型提供了一个可解的神经网络缩放定律玩具模型,同时依赖宽度与深度,并预测了随算力变化的最优Transformer形状。该模型可精确计算风险渐近值,并在源/容量条件下导出幂律关系。
原文摘要 · Abstract (English)
We study in-context learning (ICL) of linear regression in a deep linear self-attention model, characterizing how performance depends on various computational and statistical resources (width, depth, number of training steps, batch size and data per context). In a joint limit where data dimension, context length, and residual stream width scale proportionally, we analyze the limiting asymptotics for three ICL settings: (1) isotropic covariates and tasks (ISO), (2) fixed and structured covariance (FS), and (3) where covariances are randomly rotated and structured (RRS). For ISO and FS settings, we find that depth only aids ICL performance if context length is limited. Alternatively, in the RRS setting where covariances change across contexts, increasing the depth leads to significant improvements in ICL, even at infinite context length. This provides a new solvable toy model of neural scaling laws which depends on both width and depth of a transformer and predicts an optimal transformer shape as a function of compute. This toy model enables computation of exact asymptotics for the risk as well as derivation of powerlaws under source/capacity conditions for the ICL tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。