证明Transformer在动态系统建模中深度越深越优,单层效果差且依赖数据分布。
In-Context Learning of Linear Dynamical Systems with Transformers: Approximation Bounds and Depth-Separation
- 用多层Transformer逼近噪声线性动态系统,误差可接近最小二乘法。
- 单层线性Transformer的逼近误差有下界,无法随任务数量减少。
- 首次揭示Transformer在动态系统学习中的深度分离现象,适合研究者参考。
本文研究Transformer在上下文学习中对一类含噪声线性动态系统进行表示的逼近理论性质。首个理论结果建立了多层Transformer在统一任务上的$L^2$测试损失下的逼近误差上界,表明对数深度的Transformer可达到与最小二乘估计器相当的误差水平。相比之下,第二个结果为一类单层线性Transformer建立了非衰减的逼近误差下界,揭示了Transformer在动态系统上下文学习中存在深度分离现象。此外,该结果还揭示了单层线性Transformer在独立同分布(IID)与非独立同分布(non-IID)数据下逼近能力的关键差异。
原文摘要 · Abstract (English)
This paper investigates approximation-theoretic aspects of the in-context learning capability of the transformers in representing a family of noisy linear dynamical systems. Our first theoretical result establishes an upper bound on the approximation error of multi-layer transformers with respect to an $L^2$-testing loss uniformly defined across tasks. This result demonstrates that transformers with logarithmic depth can achieve error bounds comparable with those of the least-squares estimator. In contrast, our second result establishes a non-diminishing lower bound on the approximation error for a class of single-layer linear transformers, which suggests a depth-separation phenomenon for transformers in the in-context learning of dynamical systems. Moreover, this second result uncovers a critical distinction in the approximation power of single-layer linear transformers when learning from IID versus non-IID data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。