arXiv:2504.11083quant-phcs.AI2025-04被引 2

用量子退火优化注意力,实现长序列线性扩展的高效模型。

QAMA: Scalable Quantum Annealing Multi-Head Attention Operator for Deep Learning

  • 将注意力转为能量优化问题,通过量子退火自然生成稀疏结构。
  • 在多任务上精度损失≤2.7点,序列长度仅需线性量子比特。
  • 可在真实量子硬件运行,适合追求算力突破的研究者。

注意力机制是现代深度学习的核心,但其二次时间与空间复杂度限制了长序列的可扩展性。为此,本文提出量子退火多头注意力(QAMA),一种可直接替换的新型算子,将注意力重构成基于能量的哈密顿量优化问题。在此框架中,令牌间交互被编码为二元二次项,利用量子退火搜索对应有效注意力模式的低能构型。与依赖人工启发式方法的古典稀疏或近似注意力不同,QAMA让稀疏结构自然涌现。理论分析基于单自旋翻转动力学,给出依赖退火哈密顿量谱性质的求解时间边界。实验表明,在自然语言与视觉基准测试中,准确率偏差不超过2.7个百分点,且所需量子比特数随序列长度呈线性增长。可视化显示哈密顿量惩罚项诱导出有意义且可解释的跨头稀疏性。最终,在相干伊辛机上的部署验证了其在真实量子硬件上的可行性,相比经典实现展现出显著的推理时延降低。这些结果标志着将量子优化设备融入深度神经网络架构的重要一步,提供了一种无缝集成且硬件兼容的注意力替代方案。该工作已提交至IEEE,版权可能转移,后续版本可能不再开放。

原文摘要 · Abstract (English)

Attention mechanisms underpin modern deep learning, while the quadratic time and space complexity limit scalability for long sequences. To address this, Quantum Annealing Multi-Head Attention (QAMA) is proposed, a novel drop-in operator that reformulates attention as an energy-based Hamiltonian optimization problem. In this framework, token interactions are encoded into binary quadratic terms, and quantum annealing is employed to search for low-energy configurations that correspond to effective attention patterns. Unlike classical sparse or approximate attention methods that rely on hand-crafted heuristics, QAMA allows sparsity structures to emerge naturally from the optimization process. Theoretically, computational complexity is analysed through single-spin flip dynamics, providing time to solution runtime bounds that depend on the spectral properties of the annealing Hamiltonian. Empirically, evaluation on both natural language and vision benchmarks shows that, across tasks, accuracy deviates by at most 2.7 points from standard multi-head attention, while requiring only linear qubits in sequence length. Visualizations further reveal that the Hamiltonian penalty terms induce meaningful and interpretable sparsity across heads. Finally, deployment on a coherent Ising machine validates the feasibility of running QAMA on real quantum hardware, showing tangible inference-time reductions compared with classical implementations. These results highlight QAMA as a pioneering and scalable step toward integrating quantum optimization devices into deep neural architectures, providing a seamlessly integrable and hardware-compatible alternative to conventional attention mechanisms. This work has been submitted to the IEEE for possible publication. Copyright may be transferred without notice, after which this version may no longer be accessible.

量子计算注意力机制稀疏性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。