arXiv:2502.13166quant-phcs.AI2025-02ACL被引 5

用大模型动态生成量子神经网络初始参数,解决训练梯度消失问题

Large Language Models Can Help Mitigate Barren Plateaus in Quantum Neural Networks

  • 基于大模型与次鞅理论,自适应生成可避免梯度消失的初始参数
  • 在不同规模量子神经网络上,梯度方差显著高于传统初始化方法
  • 适合研究量子机器学习、特别是需应对梯度消失的算法开发者

在噪声中等规模量子(NISQ)计算时代,量子神经网络(QNNs)在诸多应用中展现出潜力,但其训练常受陡峭平原(BPs)困扰,即随着量子比特数量增加,梯度方差呈指数级衰减。现有基于初始化的缓解策略多依赖预设静态参数分布,难以适应不同模型规模或数据条件。为此,我们提出AdaInit,一种利用具备次鞅性质的大语言模型,迭代生成可保持非零梯度方差的QNN初始参数的框架。不同于传统一次性初始化,AdaInit通过融合数据特征与梯度反馈,自适应探索参数空间,并具有收敛至有效初始参数的理论保证。我们提供了次鞅过程的严格理论分析,并实证表明,AdaInit在多种QNN规模下均能持续维持更高梯度方差,显著优于现有方法。我们认为此项工作或开启缓解陡峭平原的新路径。

原文摘要 · Abstract (English)

In the era of noisy intermediate-scale quantum (NISQ) computing, Quantum Neural Networks (QNNs) have emerged as a promising approach for various applications, yet their training is often hindered by barren plateaus (BPs), where gradient variance vanishes exponentially as the qubit size increases. Most initialization-based mitigation strategies rely heavily on pre-designed static parameter distributions, thereby lacking adaptability to diverse model sizes or data conditions. To address these limitations, we propose AdaInit, a foundational framework that leverages large language models with the submartingale property to iteratively synthesize initial parameters for QNNs that yield non-negligible gradient variance, thereby mitigating BPs. Unlike conventional one-shot initialization methods, AdaInit adaptively explores the parameter space by incorporating dataset characteristics and gradient feedback, with theoretical guarantees of convergence to finding a set of effective initial parameters for QNNs. We provide rigorous theoretical analyses of the submartingale-based process and empirically validate that AdaInit consistently outperforms existing initialization methods in maintaining higher gradient variance across various QNN scales. We believe this work may initiate a new avenue to mitigate BPs.

量子机器学习梯度消失大模型应用初始化方法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。