用量子算法生成二值权重,提升低精度神经网络训练效果
Variational Inference for Quantum HyperNetworks
- 用变分量子算法通过测量生成二值权重
- 在模拟和实际场景中均优于传统最大似然训练
- 适合关注量子机器学习与模型压缩的研究者
二值神经网络(BiNNs)采用单比特精度权重,在降低内存占用和功耗的同时保持大模型的竞争力。然而,传统训练方法难以有效优化。本文提出基于量子超网络的新型训练范式,利用变分量子算法通过量子电路测量生成二值权重,借助叠加与纠缠等量子现象拓展解空间搜索范围。当可直接访问输出分布时(如仿真),我们推导出证据下界(ELBO);对于实际中常见的隐式分布,则引入基于最大均值差异(MMD)的代理ELBO。实验表明,该方法显著提升训练稳定性与泛化能力,优于标准最大似然估计(MLE)。
原文摘要 · Abstract (English)
Binary Neural Networks (BiNNs), which employ single-bit precision weights, have emerged as a promising solution to reduce memory usage and power consumption while maintaining competitive performance in large-scale systems. However, training BiNNs remains a significant challenge due to the limitations of conventional training algorithms. Quantum HyperNetworks offer a novel paradigm for enhancing the optimization of BiNN by leveraging quantum computing. Specifically, a Variational Quantum Algorithm is employed to generate binary weights through quantum circuit measurements, while key quantum phenomena such as superposition and entanglement facilitate the exploration of a broader solution space. In this work, we establish a connection between this approach and Bayesian inference by deriving the Evidence Lower Bound (ELBO), when direct access to the output distribution is available (i.e., in simulations), and introducing a surrogate ELBO based on the Maximum Mean Discrepancy (MMD) metric for scenarios involving implicit distributions, as commonly encountered in practice. Our experimental results demonstrate that the proposed methods outperform standard Maximum Likelihood Estimation (MLE), improving trainability and generalization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。