证明了Transformer可逼近光滑函数,为回归提供理论支撑
Approximation Bounds for Transformer Networks with Application to Regression
- 基于柯尔莫哥洛夫定理设计新证明思路,揭示自注意力机制的近似能力
- 在Lp范数下,误差ε所需参数量为ε^{-d_x n / γ},达最优水平
- 适用于带依赖性的非参数回归,无需限制权重大小
我们研究了Transformer网络对霍尔德与索博列夫函数的逼近能力,并将其应用于具有依赖观测的非参数回归估计。首先,建立了标准Transformer网络逼近分量函数为霍尔德连续(光滑度γ∈(0,1])的序列到序列映射的新上界。在Lp范数(p∈[1,∞])下实现误差ε,仅需固定深度的Transformer,其总参数量为ε^{-d_x n / γ}。该结果不仅将已有结论扩展至p=∞情形,且与固定深度前馈神经网络和循环神经网络的最优参数上界一致。对索博列夫函数也获得了类似结果。其次,在各种β-混合数据假设下,推导出非参数回归的显式收敛速率,允许观测间依赖随时间减弱。样本复杂度上界不施加权重幅度约束。最后,提出一种受柯尔莫哥洛夫-阿诺德表示定理启发的新证明策略,表明若Transformer的自注意力层能执行列平均,则可逼近序列到序列的霍尔德函数,为自注意力机制的可解释性提供了新视角。
原文摘要 · Abstract (English)
We explore the approximation capabilities of Transformer networks for Hölder and Sobolev functions, and apply these results to address nonparametric regression estimation with dependent observations. First, we establish novel upper bounds for standard Transformer networks approximating sequence-to-sequence mappings whose component functions are Hölder continuous with smoothness index $γ\in (0,1]$. To achieve an approximation error $\varepsilon$ under the $L^p$-norm for $p \in [1, \infty]$, it suffices to use a fixed-depth Transformer network whose total number of parameters scales as $\varepsilon^{-d_x n / γ}$. This result not only extends existing findings to include the case $p = \infty$, but also matches the best known upper bounds on number of parameters previously obtained for fixed-depth FNNs and RNNs. Similar bounds are also derived for Sobolev functions. Second, we derive explicit convergence rates for the nonparametric regression problem under various $β$-mixing data assumptions, which allow the dependence between observations to weaken over time. Our bounds on the sample complexity impose no constraints on weight magnitudes. Lastly, we propose a novel proof strategy to establish approximation bounds, inspired by the Kolmogorov-Arnold representation theorem. We show that if the self-attention layer in a Transformer can perform column averaging, the network can approximate sequence-to-sequence Hölder functions, offering new insights into the interpretability of self-attention mechanisms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。