arXiv:2602.20555stat.MLcs.IT2026-02

证明标准Transformer可逼近霍尔德函数并达到最优回归速率

Standard Transformers Achieve the Minimax Rate in Nonparametric Regression with $C^{s,λ}$ Targets

  • 用细粒度结构指标分析Transformer,建立理论基础
  • 在非参数回归中实现霍尔德函数的极小极大最优率
  • 适合关注Transformer理论性能的研究者

Transformer模型在大语言模型和计算机视觉中的巨大成功,亟需严格的理论支撑。本文首次证明,标准Transformer可在$ L^t $距离($ t ∈ [1, ∞] $)下以任意精度逼近霍尔德函数 $ C^{s,λ}([0,1]^{d imes n}) $($ s ∈ ℤ_{\geq 0}, 0 < λ ≤ 1 $)。基于此逼近结果,我们进一步证明标准Transformer在非参数回归任务中对霍尔德目标函数实现了极小极大最优率。通过引入“大小元组”与“维度向量”两个指标,本文对Transformer结构进行了细粒度刻画,为未来研究不同结构下的泛化与优化误差提供了框架。作为中间成果,还推导出标准Transformer的Lipschitz常数上界及其记忆容量,具有独立研究价值。这些发现为Transformer的强大能力提供了理论支持。

原文摘要 · Abstract (English)

The tremendous success of Transformer models in fields such as large language models and computer vision necessitates a rigorous theoretical investigation. To the best of our knowledge, this paper is the first work proving that standard Transformers can approximate Hölder functions $ C^{s,λ}\left([0,1]^{d\times n}\right) $$ (s\in\mathbb{N}_{\geq0},0<λ\leq1) $ under the $L^t$ distance ($t \in [1, \infty]$) with arbitrary precision. Building upon this approximation result, we demonstrate that standard Transformers achieve the minimax optimal rate in nonparametric regression for Hölder target functions. It is worth mentioning that, by introducing two metrics: the size tuple and the dimension vector, we provide a fine-grained characterization of Transformer structures, which facilitates future research on the generalization and optimization errors of Transformers with different structures. As intermediate results, we also derive the upper bounds for the Lipschitz constant of standard Transformers and their memorization capacity, which may be of independent interest. These findings provide theoretical justification for the powerful capabilities of Transformer models.

Transformer非参数回归理论分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。