揭示Transformer如何用注意力表达上下文关系,证明其能逼近任意关系规则。
On the Expressive Power of Contextual Relations in Transformers
- 将上下文关系建模为概率对象,统一了注意力与最优传输的关系
- 证明标准softmax注意力可逼近任意上下文关系,归因于归一化方式的选择
- 为Transformer为何有效提供了理论解释,适合研究模型原理者阅读
Transformer架构在建模上下文关系方面取得了显著的实证成功,但对其表达能力的理解仍不清晰。本文引入测度论框架,将上下文关系视为条件分布或联合分布(耦合),揭示标准softmax注意力与熵正则化最优传输之间的自然联系,统一了注意力作为底层亲和函数归一化的视角。在此框架下,我们建立了使用标准Softmax Attention与交替Sinkhorn归一化的上下文系统的通用逼近定理。结果表明,Transformer架构能够逼近任意上下文关系规则,且归一化方式决定了这些关系的表征方式。此外,该结果为Transformer在建模上下文关系方面的有效性提供了原则性解释。
原文摘要 · Abstract (English)
Transformer architectures have achieved remarkable empirical success in modeling contextual relations, yet a clear understanding of their expressive power is still lacking. In this work, we introduce a measure-theoretic framework in which contextual relations are modeled as probabilistic objects, either as conditional distributions or as joint distributions (couplings). This perspective reveals a natural connection between standard softmax attention and entropy-regularized optimal transport, providing a unified view of attention as a normalization of an underlying affinity function. Within this framework, we establish a universal approximation theorem for contextual systems using standard Softmax Attention and alternately Sinkhorn normalization. These results show that Transformer architectures can approximate arbitrary contextual relations rules, and that the choice of normalization determines how these relations are represented. Moreover, they provide a principled explanation for why Transformers are effective at modeling contextual relations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。