arXiv:2605.07705cs.LOcs.AI2026-05

用新逻辑语言揭示编码器-解码器变换器的运行机制。

Cross-Attention and Encoder-Decoder Transformers: A Logical Characterization

  • 提出一种带计数和过去时态的时序逻辑来描述模型行为。
  • 证明该逻辑能精确刻画浮点数与软注意力下的实际模型表现。
  • 适用于自回归生成等场景,对掩码等结构变化也具鲁棒性。

我们为编码器-解码器变换器(encoder-decoder transformers)提供了一种新颖的逻辑表征,这是大语言模型的基础架构,也广泛应用于依赖跨注意力机制的多种任务。我们在浮点数与软注意力的实际设置下研究此类模型,采用一种扩展命题逻辑的新时序逻辑进行刻画,该逻辑包含对编码器输入的计数全局模态和对解码器输入的过去模态。此外,我们还通过一类分布式自动机给出了另一种表征,并证明结果不局限于架构中的特定选择,可适应如掩码等结构变化。最后,我们讨论了在自回归生成设定下的编码器-解码器变换器。

原文摘要 · Abstract (English)

We give a novel logical characterization of encoder-decoder transformers, the foundational architecture for LLMs that also sees use in various settings that benefit from cross-attention. We study such transformers over text in the practical setting of floating-point numbers and soft-attention, characterizing them with a new temporal logic. This logic extends propositional logic with a counting global modality over the encoder input and a past modality over the decoder input. We also give an additional characterization of such transformers via a type of distributed automata, and show that our results are not limited to the specific choices in the architecture and can account for changes in, e.g., masking. Finally, we discuss encoder-decoder transformers in the autoregressive setting.

变换器逻辑表征自回归

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。