揭示Transformer通用逼近的理论条件,为新型架构设计提供保障。
A unified framework for establishing the universal approximation of transformer-type architectures
- 提出统一理论框架,以令牌可区分性为核心条件
- 证明多种注意力机制的Transformer均具通用逼近能力
- 适用于新架构设计,尤其适合有对称性约束的场景
本文研究Transformer类架构的通用逼近性质(UAP),构建了一个统一的理论框架,将先前针对残差网络的结果扩展至包含注意力机制的模型。研究发现,令牌可区分性是实现UAP的基本要求,并提出了一个适用于广泛架构的一般充分条件。在注意力层解析性假设下,该条件的验证可显著简化,从而非构造性地建立了此类架构的UAP。通过该框架,我们证明了采用多种注意力机制(包括基于核和稀疏注意力)的Transformer均具备通用逼近能力。结果既推广了已有工作,也首次为此前未覆盖的架构建立了UAP。此外,该框架为设计具有内在UAP保证的新Transformer架构提供了原则性基础,包括具有特定函数对称性的结构,并给出了具体示例。
原文摘要 · Abstract (English)
We investigate the universal approximation property (UAP) of transformer-type architectures, providing a unified theoretical framework that extends prior results on residual networks to models incorporating attention mechanisms. Our work identifies token distinguishability as a fundamental requirement for UAP and introduces a general sufficient condition that applies to a broad class of architectures. Leveraging an analyticity assumption on the attention layer, we can significantly simplify the verification of this condition, providing a non-constructive approach in establishing UAP for such architectures. We demonstrate the applicability of our framework by proving UAP for transformers with various attention mechanisms, including kernel-based and sparse attention mechanisms. The corollaries of our results either generalize prior works or establish UAP for architectures not previously covered. Furthermore, our framework offers a principled foundation for designing novel transformer architectures with inherent UAP guarantees, including those with specific functional symmetries. We propose examples to illustrate these insights.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。