量子注意力机制用少量参数实现高阶令牌交互,突破经典模型限制。
Higher-Order Token Interactions via Quantum Attention

- 通过量子电路数据重加载与非克里福纠缠,单层实现任意阶交互。
- 参数量仅经典模型的1/6.5却能泛化至6阶隐藏子集奇偶性任务。
- 适合需要高阶特征建模的领域,如基因互作、噪声奇偶学习等。
标准点积自注意力在单层中仅能计算令牌间的成对(二阶)交互;表示通用的k阶交互通常需超二次资源或多层堆叠。本文提出量子高阶注意力(QHA),一种浅层且可硬件实现的量子注意力头,通过数据重加载和全连接非克里福纠缠,在电路内合成k阶令牌交互,并通过局部单量子比特读出暴露。理论证明:(i) 任何嵌入维数m、H个头、p位精度满足mHp=o(N/log log N)的标准自注意力层,无法表示一个单个QHA头以深度O(log k)(O(k)两量子比特门)所表示的k阶相关性族;(ii) 其局部设计实例具有可训练性保障:局部读出且深度O(log n)下梯度方差为Ω(1/poly(n))(无消失梯度),实证验证成立——而全连接实例训练时出现指数级梯度衰减。实验显示,在6.5倍更小的参数预算下,QHA能从不相交输入中泛化所有k≤6阶隐藏子集奇偶性任务,而更大规模的经典注意力头在超过二阶后即崩溃;优势大小与目标傅里叶度正相关,奇偶性最大,低阶结构存在时减弱。应用上,QHA作为紧凑高阶交互探测器,在遗传上位性、带噪声奇偶学习、图三角检测三个领域均达到噪声极限,且在经典线性方法失效的最小参数量下表现最优。
原文摘要 · Abstract (English)
Standard dot-product self-attention computes, in a single layer, only pairwise (order-2) interactions between tokens; representing a generic order-$k$ interaction is known to require either super-quadratic resources in one layer or composition across depth. We introduce \textbf{Quantum Higher-Order Attention (QHA)}, a shallow, hardware-realizable quantum attention head that, via data re-uploading and an all-to-all non-Clifford entangler, synthesizes order-$k$ token interactions inside the circuit and exposes them through a local single-qubit read-out. We prove (i) an expressivity separation: any single standard self-attention layer with embedding dimension $m$, $H$ heads and $p$-bit precision satisfying $mHp=o(N/\log\log N)$ cannot represent the order-$k$ correlation family that one QHA head represents with circuit depth $O(\log k)$ ($O(k)$ two-qubit gates); and (ii) a trainability guarantee for its local-design instantiation: with a local read-out and $O(\log n)$ depth the gradient variance is $Ω(1/\mathrm{poly}(n))$ (no barren plateau), which we confirm empirically -- while being explicit that the more expressive all-to-all instantiation we benchmark is trained empirically and shows exponentially decaying gradients. Empirically, at a $6.5\times$ smaller parameter budget, QHA generalizes hidden-subset parity of every order $k\le6$ from disjoint inputs, whereas the larger classical attention head collapses past order~2; consistent with theory, the size of the advantage tracks the target's Fourier degree - largest for parity and shrinking when low-order structure is present. As an application, QHA serves as a compact high-order interaction detector across three domains - genetic epistasis, learning-parity-with-noise, and graph triangle detection - reaching the noise ceiling at the smallest parameter budget where field-standard linear methods fail.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。