用分布回归理论分析Transformer泛化能力,揭示其函数学习优势。
Generalization Analysis of Transformers in Distribution Regression
- 构建基于分布回归的Transformer框架,定义注意力算子实现无损分布压缩。
- 证明Transformer在复杂函数建模上优于CNN与全连接网络。
- 为提示调优等大模型技术提供新理论解释,适合研究者参考。
近年来,基于Transformer架构的模型广泛应用并成为深度学习核心工具。尽管参数高效微调和高效扩展等技术取得成功,但缺乏严格的数学理论支持。本文提出一种受分布回归启发的Transformer学习框架,将分布作为输入,连接两阶段采样过程与自然语言处理,并提出注意力算子的数学形式。证明通过注意力算子,Transformer可无损地将分布压缩为函数表示。同时,得益于该算子优势,Transformer在学习复杂结构函数时表现出强于卷积神经网络和全连接网络的能力。最后,在分布回归框架下获得泛化界。基于上述理论结果,进一步探讨大语言模型中出现的成功技术,如提示调优、参数高效微调和高效扩展,并在新分析框架下提供理论洞察。
原文摘要 · Abstract (English)
In recent years, models based on the Transformer architecture have seen widespread applications and have become one of the core tools in the field of deep learning. Numerous successful techniques, such as parameter-efficient fine-tuning and efficient scaling, have been proposed surrounding their applications to further enhance performance. However, the success of these strategies has always lacked the support of rigorous mathematical theory. To study the underlying mechanisms behind Transformers and related techniques, we first propose a Transformer learning framework motivated by distribution regression, with distributions being inputs, connect a two-stage sampling process with natural language processing, and present a mathematical formulation of the attention mechanism called attention operator. We demonstrate that by the attention operator, Transformers can compress distributions into function representations without loss of information. Moreover, with the advantages of our novel attention operator, Transformers exhibit a stronger capability to learn functionals with more complex structures than convolutional neural networks and fully connected networks. Finally, we obtain a generalization bound within the distribution regression framework. Through the aforementioned theoretical results, we further discuss some successful techniques emerging with large language models (LLMs), such as prompt tuning, parameter-efficient fine-tuning, and efficient scaling. We also provide theoretical insights behind these techniques within our novel analysis framework.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。