arXiv:2606.12478cs.LGcond-mat.stat-mech2026-06

用可学习的相互作用建模注意力协同,提升长序列建模能力

Boltzmann Attention: Learnable Ising Couplings for Cooperative Attention

论文配图:Boltzmann Attention: Learnable Ising Couplings for Cooperative Attention
图 1 · 摘自论文原文
  • 引入基于伊辛模型的能量机制,显式建模注意力位置间协同关系
  • 在字符级语言建模和括号匹配任务中,长序列下性能优于标准注意力
  • 支持量子退火训练,为未来量子计算应用提供路径

注意力机制是现代序列模型的核心,但标准注意力主要依赖查询与键的个体相似性。尽管软最大归一化引入了位置间的竞争,但标准注意力层并未显式参数化注意力决策间的可学习交互,限制了其直接建模注意力结构中的协同或对抗关系的能力。本文提出玻尔兹曼注意力,一种基于能量的推广方法,其中注意力模式由相互作用的伊辛模型决定。该方法在依赖数据的局部场基础上增加可学习的成对耦合,使模型能够表示超出软最大或符号函数注意力所捕捉的跨位置相关性。在字符级语言建模和合成括号匹配任务上的实验表明,玻尔兹曼注意力在标准Transformer架构中始终优于标准软最大注意力,且随着序列长度增加,优势更加显著。四组消融实验证实性能提升源于可学习的成对耦合。结果表明,显式的位置间相互作用为注意力序列建模提供了原则性的增强。此外,伊辛形式自然引出基于量子计算的采样策略:我们证明绝热量子退火是一种可行的训练方法,同时保持与精确玻尔兹曼计算相当的性能。

原文摘要 · Abstract (English)

Attention mechanisms are central to modern sequence models, yet standard attention computes relevance primarily through individual query--key similarities. Although softmax normalization introduces competition among positions, a standard attention layer does not explicitly parameterize learnable interactions between attention decisions. This limits its ability to directly model cooperative or antagonistic co-attention structure within the attention mechanism itself. We propose Boltzmann attention, an energy-based generalization in which attention patterns are governed by an interacting Ising model. The method augments the usual data-dependent local fields with learnable pairwise couplings, allowing the model to represent inter-position correlations beyond those captured by softmax or sigmoid attention. Experiments on character-level language modeling and synthetic bracket matching show that Boltzmann attention consistently improves over standard softmax attention within a standard Transformer architecture, with the advantage becoming more pronounced as sequence length increases. A four-way ablation confirms that the improvement arises from the learnable pairwise couplings. These results suggest that explicit inter-position interactions provide a principled enhancement for attention-based sequence modeling. Moreover, the Ising formulation opens a natural path toward quantum-computing-based sampling strategies: we demonstrate that diabatic quantum annealing provides a practical training method while maintaining competitive performance with exact Boltzmann computation.

注意力机制伊辛模型量子计算序列建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。