研究量子神经网络的局部参数空间可区分性,发现小模型难训练
Exploring Channel Distinguishability in Local Neighborhoods of the Model Space in Quantum Neural Networks
- 从参数邻域分析量子电路的可区分性,提出新评估视角
- 证明小参数模型在微小扰动下几乎不可区分,难以优化
- 提示需精心初始化,不适用于迭代训练小模型
随着量子机器学习兴起,量子神经网络(QNN)受到广泛关注。然而,这些模型极难训练,我们推测部分原因在于其架构(ansatz)尚未被充分研究。本文回溯分析ansatz,首先考察其表达能力——即能表示的算子空间,发现主流的2设计接近度无法有效刻画此属性。因此,我们转向模型空间的局部邻域,通过分析参数微小扰动下的模型可区分性来表征ansatz。我们推导出可区分性的上界,表明参数少的QNN在参数更新后几乎无法区分。数值实验支持该结论,并揭示显著的差异性,强调了热启动或智能初始化的重要性。整体而言,本工作从ansatz视角揭示了QNN训练动态与困难,暗示小量子模型的迭代训练可能无效,与其初始动机相悖。
原文摘要 · Abstract (English)
With the increasing interest in Quantum Machine Learning, Quantum Neural Networks (QNNs) have emerged and gained significant attention. These models have, however, been shown to be notoriously difficult to train, which we hypothesize is partially due to the architectures, called ansatzes, that are hardly studied at this point. Therefore, in this paper, we take a step back and analyze ansatzes. We initially consider their expressivity, i.e., the space of operations they are able to express, and show that the closeness to being a 2-design, the primarily used measure, fails at capturing this property. Hence, we look for alternative ways to characterize ansatzes by considering the local neighborhood of the model space, in particular, analyzing model distinguishability upon small perturbation of parameters. We derive an upper bound on their distinguishability, showcasing that QNNs with few parameters are hardly discriminable upon update. Our numerical experiments support our bounds and further indicate that there is a significant degree of variability, which stresses the need for warm-starting or clever initialization. Altogether, our work provides an ansatz-centric perspective on training dynamics and difficulties in QNNs, ultimately suggesting that iterative training of small quantum models may not be effective, which contrasts their initial motivation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。