arXiv:2602.11042quant-phcs.LG2026-02被引 9

研究量子生成模型的可训练性,发现特定参数初始化能避免梯度消失。

Characterizing Trainability of Instantaneous Quantum Polynomial Circuit Born Machines

  • 通过解析推导损失梯度方差,揭示梯度消失与生成器结构和核谱的关系。
  • 低权重偏置核在结构化拓扑中可避免指数级梯度衰减,保持可训练性。
  • 小方差高斯初始化下梯度可多项式缩放,适合低频成分的训练。

瞬时量子多项式电路玻恩机(IQP-QCBMs)作为一类具有经典可处理训练目标的量子生成模型,其最大均值差异(MMD)目标函数和采样复杂性论证暗示了潜在的量子优势,使其成为值得深入研究的模型。尽管已有研究证明其扩展模型的通用性,但当前亟需探讨其可训练性:是否受指数级消失梯度(即贫瘠高原问题)影响,从而阻碍有效训练,并且可训练区域与量子优势区域是否存在重叠。本文在此方向取得重要进展:针对训练初始状态的可训练性,我们解析推导了MMD损失函数偏导数方差的闭式表达式,给出一般上下界;在均匀初始化下,表明贫瘠高原依赖于生成器集合和所选核的谱特性;识别出低权重偏置核在结构化拓扑中可避免指数梯度抑制的区域。此外,我们证明在温和条件下,小方差高斯初始化可保证梯度的多项式尺度。关于潜在量子优势,基于先前复杂性理论论证,我们进一步指出稀疏的IQP族可输出经典难以模拟的概率分布族,且该分布至少在低权重频率下可在初始化时保持可训练。

原文摘要 · Abstract (English)

Instantaneous quantum polynomial quantum circuit Born machines (IQP-QCBMs) have been proposed as quantum generative models with a classically tractable training objective based on the maximum mean discrepancy (MMD) and a potential quantum advantage motivated by sampling-complexity arguments, making them an exciting model worth deeper investigation. While recent works have further proven the universality of a (slightly generalized) model, the next immediate question pertains to its trainability, i.e., whether it suffers from the exponentially vanishing loss gradients, known as the barren plateau issue, preventing effective use, and how regimes of trainability overlap with regimes of possible quantum advantage. Here, we provide significant strides in these directions. To study the trainability at initialization, we analytically derive closed-form expressions for the variances of the partial derivatives of the MMD loss function and provide general upper and lower bounds. With uniform initialization, we show that barren plateaus depend on the generator set and the spectrum of the chosen kernel. We identify regimes in which low-weight-biased kernels avoid exponential gradient suppression in structured topologies. Also, we prove that a small-variance Gaussian initialization ensures polynomial scaling for the gradient under mild conditions. As for the potential quantum advantage, we further argue, based on previous complexity-theoretic arguments, that sparse IQP families can output a probability distribution family that is classically intractable, and that this distribution remains trainable at initialization at least at lower-weight frequencies.

量子生成模型可训练性贫瘠高原量子优势

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。