提出基于矩阵秩的Transformer泛化误差新边界,揭示低秩设计优势。
On Rank-Dependent Generalisation Error Bounds for Transformers
- 基于矩阵秩构建覆盖数边界,用于分析单层Transformer性能。
- 泛化误差随样本量呈1/√n衰减,优于现有文献的(log n)/√n结果。
- 适用于关注模型泛化性与低秩结构设计的研究者。
本文为满足不同输入与矩阵范数约束的线性函数类引入多种覆盖数界,这些界依赖于矩阵类的秩。将这些界应用于单层Transformer,推导出其泛化误差界。结果改进了文献中多项已有泛化界,且不依赖输入序列长度,凸显了低秩矩阵在Transformer设计中的优势。具体而言,所得泛化误差界以O(1/√n)速率衰减(n为样本量),优于现有研究中O((log n)/√n)的阶;同时以O(log r_w)速率随查询与键矩阵组合的秩r_w衰减。
原文摘要 · Abstract (English)
In this paper, we introduce various covering number bounds for linear function classes, each subject to different constraints on input and matrix norms. These bounds are contingent on the rank of each class of matrices. We then apply these bounds to derive generalization errors for single layer transformers. Our results improve upon several existing generalization bounds in the literature and are independent of input sequence length, highlighting the advantages of employing low-rank matrices in transformer design. More specifically, our achieved generalisation error bound decays as $O(1/\sqrt{n})$ where $n$ is the sample length, which improves existing results in research literature of the order $O((\log n)/(\sqrt{n}))$. It also decays as $O(\log r_w)$ where $r_w$ is the rank of the combination of query and and key matrices.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。