arXiv:2509.00421cs.LG2025-09被引 3

揭示提示调优在长文本中的记忆瓶颈,解释为何大模型处理长上下文会性能下降。

Memory Limitations of Prompt Tuning in Transformers

  • 证明提示长度增长时,模型记忆信息量最多线性增加
  • 首次严格证明长上下文导致性能下降的内在原因
  • 适用于研究模型长期依赖与提示工程的学者

尽管提示调优在适应预训练语言模型至新任务上表现出色,但其理论分析仍有限。现有理论工作主要关注通用近似能力,结果与标准权重调优相当。本文从另一角度探讨Transformer的理论:提示调优的记忆能力。我们提出两项主要理论贡献:首先,证明Transformer所记忆的信息量随提示长度的增长不可能快于线性;其次,也是更重要的一点,我们首次给出了对大语言模型中观测到的现象的正式证明——长上下文下的性能退化。我们严格证明,Transformer具有内在的记忆限制,无论上下文多长,能保留的信息量都有上限。这一发现为Transformer架构在处理长序列时的能力局限提供了根本性理解。

原文摘要 · Abstract (English)

Despite the empirical success of prompt tuning in adapting pretrained language models to new tasks, theoretical analyses of its capabilities remain limited. Existing theoretical work primarily addresses universal approximation properties, demonstrating results comparable to standard weight tuning. In this paper, we explore a different aspect of the theory of transformers: the memorization capability of prompt tuning. We provide two principal theoretical contributions. First, we prove that the amount of information memorized by a transformer cannot scale faster than linearly with the prompt length. Second, and more importantly, we present the first formal proof of a phenomenon empirically observed in large language models: performance degradation in transformers with extended contexts. We rigorously demonstrate that transformers inherently have limited memory, constraining the amount of information they can retain, regardless of the context size. This finding offers a fundamental understanding of the intrinsic limitations of transformer architectures, particularly their ability to handle long sequences.

提示调优记忆瓶颈Transformer

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。