arXiv:2511.01918cs.LGquant-ph2025-11中稿 · 2025 IEEE Internat…

用量子叠加原理改进训练,让模型收敛更快更准。

Superpositional Gradient Descent: Harnessing Quantum Principles for Model Training

  • 将量子电路扰动注入梯度更新,实现量子叠加式优化
  • 在序列分类和大模型微调中,收敛速度更快、损失更低
  • 适合对训练效率和模型性能有高要求的研究者

大型语言模型(LLMs)通常使用经典优化方法如AdamW进行训练,以提升收敛性和泛化能力。然而,量子启发方法如何增强经典训练的机制仍不明确。本文提出超位置梯度下降(Superpositional Gradient Descent, SGD),通过引入量子电路扰动,将梯度更新与量子叠加原理相联系。我们构建了数学框架,并在PyTorch与Qiskit中实现了混合量子-经典电路。在合成序列分类与大规模LLM微调任务中,SGD相比AdamW收敛更快,最终损失更低。尽管结果令人鼓舞,但可扩展性与硬件限制制约了其实际应用。本工作为量子计算与深度学习的交叉提供了新见解,揭示了利用量子原理调控和增强模型行为的可行路径。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly trained with classical optimization techniques like AdamW to improve convergence and generalization. However, the mechanisms by which quantum-inspired methods enhance classical training remain underexplored. We introduce Superpositional Gradient Descent (SGD), a novel optimizer linking gradient updates with quantum superposition by injecting quantum circuit perturbations. We present a mathematical framework and implement hybrid quantum-classical circuits in PyTorch and Qiskit. On synthetic sequence classification and large-scale LLM fine-tuning, SGD converges faster and yields lower final loss than AdamW. Despite promising results, scalability and hardware constraints limit adoption. Overall, this work provides new insights into the intersection of quantum computing and deep learning, suggesting practical pathways for leveraging quantum principles to control and enhance model behavior.

优化器量子计算大模型训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。