arXiv:2503.07294cs.CVcs.AI2025-03

量子自注意力让视觉模型参数减少99.99%,仍保持高精度。

From $\mathcal{O}(n^{2})$ to $\mathcal{O}(n)$ Parameters: Quantum Self-Attention in Vision Transformers for Biomedical Image Classification

  • 用量子神经网络替代传统自注意力,参数量从O(n²)降至O(n)
  • 在RetinaMNIST上达56.5%准确率,仅比顶尖模型低0.88%
  • 首次验证经典到量子知识蒸馏有效,适合资源受限的生物医学图像分析

我们证明,将视觉变换器(ViT)中的自注意力机制替换为量子自注意力(QSA)的量子视觉变换器(QViT),可在使用99.99%更少参数的情况下达到最先进(SOTA)的生物医学图像分类性能。QSA通过用参数化量子神经网络(QNNs)替代线性自注意力层实现,使参数量级从O(n²)降低至O(n)。在RetinaMNIST数据集上,我们的超参数高效QViT优于13/14个SOTA方法(包括CNN和ViT),达到56.5%准确率,仅比顶级MedMamba模型低0.88%,同时参数量仅为1K(对比14.5M),计算量减少89%。我们首次研究了经典到量子视觉变换器的知识蒸馏(KD),结果显示,在八个涵盖多种模态的多样化数据集上,QViT性能可与经典ViT相当,且量子自注意力参数效率更高。更高量子比特架构在预训练时受益更多,表明QSA参数量与KD有效性之间存在正相关。这些发现确立了QSA作为参数高效生物医学图像分析的实际架构选择。

原文摘要 · Abstract (English)

We demonstrate that quantum vision transformers (QViTs), vision transformers (ViTs) with self-attention (SA) mechanisms replaced by quantum self-attention (QSA) mechanisms, can match state-of-the-art (SOTA) biomedical image classifiers while using 99.99% fewer parameters. QSAs are produced by replacing linear SA layers with parameterised quantum neural networks (QNNs), producing a QSA mechanism and reducing parameter scaling from $\mathcal{O}(n^2)$ to $\mathcal{O}(n)$. On RetinaMNIST, our ultra parameter-efficient QViT outperforms 13/14 SOTA methods including CNNs and ViTs, achieving 56.5% accuracy, just 0.88% below the top MedMamba model while using 99.99% fewer parameters (1K vs 14.5M) and 89% fewer GFLOPs. We present the first investigation of knowledge distillation (KD) from classical to quantum vision transformers in biomedical image classification, showing that QViTs maintain comparable performance to classical ViTs across eight diverse datasets spanning multiple modalities, with improved QSA parameter-efficiency. Our higher-qubit architecture benefitted more from KD pre-training, suggesting a scaling relationship between QSA parameters and KD effectiveness. These findings establish QSA as a practical architectural choice toward parameter-efficient biomedical image analysis.

量子计算视觉模型参数效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。