arXiv:2411.16525cs.LGcs.AI2024-11中稿 · ICLR被引 21

揭示单头单层Transformer提示调优的理论极限

Fundamental Limits of Prompt Tuning Transformers: Universality, Capacity and Efficiency

  • 在最简Transformer上证明提示调优具通用逼近能力
  • 发现提示长度需指数级增长才能记忆数据集
  • 给出高效推理的条件,指导实际方法设计

我们研究了基于Transformer的基础模型提示调优的统计与计算极限。核心贡献在于对仅含单头、单自注意力层的Transformer进行提示调优:(i) 具有通用性,(ii) 在强指数时间假设(SETH)下支持高效(甚至近线性时间)算法。统计上,证明此类最简Transformer的提示调优是序列到序列Lipschitz函数的通用逼近器。此外,我们给出了在1层、1头Transformer上,提示调优要记忆任意数据集所需软提示标记数量的指数下界,该下界随$dL$和$(1/ε)$指数增长。计算上,我们识别出提示调优效率的相变点,由软提示引发的键和查询范数决定,并提供一个上界准则。超出此准则时,在SETH下不存在亚二次(高效)的提示调优算法;在此准则内,我们证明了近线性时间提示调优推理算法的存在性。这些基本限制为实践者设计表达能力强且高效的提示调优方法提供了重要必要条件。

原文摘要 · Abstract (English)

We investigate the statistical and computational limits of prompt tuning for transformer-based foundation models. Our key contributions are prompt tuning on \emph{single-head} transformers with only a \emph{single} self-attention layer: (i) is universal, and (ii) supports efficient (even almost-linear time) algorithms under the Strong Exponential Time Hypothesis (SETH). Statistically, we prove that prompt tuning on such simplest possible transformers are universal approximators for sequence-to-sequence Lipschitz functions. In addition, we provide an exponential-in-$dL$ and -in-$(1/ε)$ lower bound on the required soft-prompt tokens for prompt tuning to memorize any dataset with 1-layer, 1-head transformers. Computationally, we identify a phase transition in the efficiency of prompt tuning, determined by the norm of the \emph{soft-prompt-induced} keys and queries, and provide an upper bound criterion. Beyond this criterion, no sub-quadratic (efficient) algorithm for prompt tuning exists under SETH. Within this criterion, we showcase our theory by proving the existence of almost-linear time prompt tuning inference algorithms. These fundamental limits provide important necessary conditions for designing expressive and efficient prompt tuning methods for practitioners.

提示调优Transformer理论分析计算效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。