通过分块处理改善量子神经网络优化难题,提升训练稳定性。
The effect of the number of parameters and the number of local feature patches on loss landscapes in distributed quantum neural networks
- 将输入数据切分为重叠局部块,由多个独立量子网络并行处理
- 增加局部块数量可显著降低损失曲面最大赫森特征值,改善优化
- 该方法隐式引入结构正则化,适合构建可扩展的量子机器学习模型
量子神经网络有望解决经典计算机难以处理的计算难题,但其实际应用受限于复杂的损失曲面,表现为荒原平原和大量局部极小值。随着参数量或量子比特数增加,优化难度加剧。为缓解这一问题,特别是针对经典数据,本文采用分布式策略:将输入数据划分为重叠的局部块,每个块由独立的量子神经网络处理,并聚合输出进行预测。我们通过理论与实验的赫森分析及损失曲面可视化,研究了参数量和局部块数量对损失景观几何的影响。结果表明,增加参数量会加剧损失曲面的深度与陡峭度;而增加局部块数量能显著降低极小值处的最大赫森特征值。全赫森特征谱分析显示,存在大量接近零的特征值,以及与类别数对应的明显异常峰值,类似经典深度学习模型。这些发现表明,该分布式分块方法起到了隐式结构正则化作用,促进优化稳定性并可能提升泛化能力。本研究为量子机器学习在经典数据任务中的可训练性与可扩展性提供了重要见解。
原文摘要 · Abstract (English)
Quantum neural networks hold promise for tackling computationally challenging tasks that are intractable for classical computers. However, their practical application is hindered by significant optimization challenges, arising from complex loss landscapes characterized by barren plateaus and numerous local minima. These problems become more severe as the number of parameters or qubits increases, hampering effective training. To mitigate these optimization challenges, particularly for classical data, we distribute overlapping local patches across multiple quantum neural networks, processing each patch with an independent quantum neural network, and aggregating their outputs for prediction. In this study, we investigate how the number of parameters and patches affects the loss landscape geometry of this distributed quantum neural network architecture via theoretical and empirical Hessian analyses and loss landscape visualization. Our results confirm that increasing the number of parameters tends to lead to deeper and sharper loss landscapes. Crucially, we theoretically derive and empirically demonstrate that increasing the number of patches significantly reduces the largest Hessian eigenvalue at minima. Furthermore, our analysis of the full Hessian eigenspectrum reveals a structure consisting of a bulk of near-zero eigenvalues and distinct outlier spikes corresponding to the number of classes, similar to classical deep learning models. These findings suggest that our distributed patch approach acts as a form of implicit structural regularization, promoting optimization stability and potentially enhancing generalization. Our study provides valuable insights into optimization challenges and highlights that the distributed patch approach is a promising strategy for developing more trainable and scalable quantum machine learning models for classical data tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。