arXiv:2602.06699quant-phcs.CL2026-02

用量子叠加干涉实现注意力机制,可直接输出损失值并处理量子数据序列。

Quantum Attention by Overlap Interference: Predicting Sequences from Classical and Many-Body Quantum Data

  • 通过量子态重叠干涉实现非线性注意力,无需解码幅度编码结果。
  • 门复杂度为 O(T d²),在序列长度大时优于经典方法的 O(T² d)。
  • 适用于量子动力学建模,能学习经典与多体量子数据的序列规律。

我们提出一种变分量子自注意力(QSA)实现方式,作为变换器和大语言模型的核心操作,通过形成过去数据的重叠加权组合来预测序列未来元素。与以往方法不同,QSA利用量子态重叠的干涉实现所需非线性,并将瑞尼-1/2交叉熵损失直接作为可观测量的期望值输出,避免了将振幅编码的预测解码为经典逻辑值的步骤。此外,QSA自然支持可训练的数据嵌入,使量子态重叠与数据级相似性相联系。我们发现QSA的门复杂度主导项为O(T d²),而经典方法为O(T² d),表明在序列长度T远大于嵌入维度d的实际场景下具有优势。模拟结果显示,基于QSA的量子变换器能够学习经典数据和多体横场伊辛模型量子轨迹的序列预测任务,验证了可训练注意力作为量子动力学建模的实用基本单元。

原文摘要 · Abstract (English)

We propose a variational quantum implementation of self-attention (QSA), the core operation in transformers and large language models, which predicts future elements of a sequence by forming overlap-weighted combinations of past data. At variance with previous approaches, our QSA realizes the required nonlinearity through interference of state overlaps and returns a Renyi-1/2 cross-entropy loss directly as the expectation value of an observable, avoiding the need to decode amplitude-encoded predictions into classical logits. Furthermore, QSA naturally accommodates a constrained, trainable data-embedding that ties quantum state overlaps to data-level similarities. We find a gate complexity dominant scaling O(T d^2) for QSA, versus O(T^2 d) classically, suggesting an advantage in the practical regime where the sequence length T dominates the embedding size d. In simulations, we show that our QSA-based quantum transformer learns sequence prediction on classical data and on many-body transverse-field Ising quantum trajectories, establishing trainable attention as a practical primitive for quantum dynamical modeling.

量子计算注意力机制序列建模量子动力学

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。