量子注意力机制的性能提升源于架构设计,而非量子特性。
Quantum Adaptive Self-Attention for Quantum Transformer Models
- 用参数量匹配的经典瓶颈对比量子电路,验证性能来源
- 在9个合成任务和真实数据上,量子层表现与经典瓶颈相当
- 性能优势来自低秩压缩的架构原则,适合研究量子模型可解释性
量子机器学习中所谓的‘量子优势’常未与参数量匹配的经典模型对比,难以判断收益来自量子硬件还是架构改进。本文提出一种诚实归因的方法:使用相同参数预算的容量匹配经典瓶颈、透明报告量子不占优的情况,并在真实量子硬件上验证。以量子自适应注意力(QASA)为例,其仅将单个编码器层的值投影替换为36参数的量子电路,其余层保持经典。在九个合成基准和真实世界ETTh1数据集上,QASA在混沌与趋势主导信号上优于全容量经典Transformer。但引入容量匹配的经典瓶颈后,其误差指标与量子电路持平。因此,性能提升源于低秩值投影的架构简约原则,而非量子性;增加更多量子层反而降低性能与可训练性。量子层应被视为该原则的竞争力实现,而非精度优势的来源。
原文摘要 · Abstract (English)
A recurring weakness in quantum machine learning (QML) is that reported ``quantum advantages'' are seldom tested against a \emph{capacity-matched} classical control, leaving it unclear whether a gain comes from the quantum substrate or from the architectural change that accompanies it. Our primary contribution is methodological: a protocol for attributing such gains honestly -- a capacity-matched classical bottleneck of identical parameter budget, transparent reporting of where quantum does \emph{not} help, and validation on real quantum hardware -- which we develop and apply through a concrete case study. That case study is Quantum Adaptive Self-Attention (QASA), a hybrid Transformer that replaces the value projection of a \emph{single} encoder layer with a 36-parameter parameterized quantum circuit (PQC), keeping all other layers classical. Across nine synthetic benchmarks and the real-world ETTh1 dataset, QASA improves on a full-capacity classical Transformer for chaotic and trend-dominated signals. To ask whether this is a genuinely \emph{quantum} effect, we introduce a control rarely applied in quantum machine learning -- a capacity-matched classical bottleneck with the same parameter budget -- and find that it matches the PQC on the error metrics. The gain is therefore attributable to the low-rank value-projection \emph{bottleneck} (an \emph{architectural parsimony} principle), not to quantumness; adding further quantum layers only degrades performance and trainability. We accordingly position the quantum layer not as a source of accuracy advantage but as a \emph{competitive} instantiation of this principle: its low-rank compression onto the signal's intrinsic dimensionality is matched by a classical bottleneck, so the gain is architectural rather than quantum.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。