揭示自回归Transformer生成语言时能表达的概率分布特性
Probability Distributions Computed by Autoregressive Transformers
- 从语言识别转向生成视角,分析Transformer的表达能力
- 自回归化可提升部分场景下的表达能力,概率性破坏非概率等价关系
- 为实际语言模型应用提供理论支撑,适合关注模型表达力的研究者
多数关于Transformer的表达能力研究将其视为语言识别器——判断字符串是否合法的装置,而非实际使用中的自回归概率语言模型。本文刻画了变压器语言模型所能表达的概率分布。研究表明,将语言识别器变为自回归模式有时能提升其表达能力,而引入概率机制则可能打破非概率情形下的等价关系。总体贡献在于厘清了在最常见的语言模型应用场景下,Transformer能够表达的函数类型。
原文摘要 · Abstract (English)
Most expressivity results for transformers treat them as language recognizers -- devices that accept or reject strings -- rather than as they are used in practice: as language models that generate strings autoregressively and probabilistically. We characterize the probability distributions that transformer language models can express. We show that making transformer language recognizers autoregressive can sometimes increase their expressivity, and that making them probabilistic can break equivalences that hold in the non-probabilistic case. Our overall contribution is to tease apart what functions transformers are capable of expressing in their most common use case as language models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。