arXiv:2606.21974quant-phcs.AI2026-06

用量子电路模拟训练大模型,让其真正理解量子计算逻辑。

Fine-Tuning Large Language Models for Quantum Reasoning

论文配图:Fine-Tuning Large Language Models for Quantum Reasoning
图 1 · 摘自论文原文
  • 通过逐门状态演化轨迹进行监督微调,引导模型学习量子推理。
  • 微调后模型在小系统内精度接近完美,并能外推到更大门数系统。
  • 新方法使大模型具备跨规模泛化能力,适合量子算法研究者使用。

大型语言模型(LLMs)展现出超越自然语言建模的能力,近期推理能力的提升使其适用于需要深度领域知识和复杂推理的科学任务。量子计算作为高度专业化领域,受限于知识门槛与硬件条件,可从该进展中获益。但关键问题是:如何构建微调流程,使模型具备真正的量子推理能力,而非仅依赖任务特定模式匹配?本文以量子电路模拟为训练目标,要求模型预测一系列量子门操作后的测量概率分布。提出并比较两种微调方案:(1) 在显式逐门状态矢量模拟轨迹上进行监督微调(SFT);(2) 两阶段SFT+组相对策略优化(GRPO),先进行SFT再用可验证奖励进行GRPO。结果表明,SFT在分布内及门数外推上达到近乎完美的准确率,显著优于基础模型与GPT-OSS-120B基线。SFT+GRPO虽牺牲部分分布内精度,但在处理更大量子比特系统时表现更优,远超基线。两项方法均证明,针对显式推理轨迹的定向微调是提升大模型量子推理能力的有效策略。

原文摘要 · Abstract (English)

Large language models (LLMs) exhibit abilities beyond natural language modelling and text generation. Recent advances in their reasoning capabilities have spurred interest in applying LLMs to complex scientific tasks requiring deep domain expertise and sophisticated reasoning. Quantum computing, as a highly specialised field with significant knowledge barriers and hardware constraints, could greatly benefit from such advancements. However, a key open question that first must be answered is: How can we develop fine-tuning pipelines that instil genuine quantum reasoning in LLMs, rather than task-specific pattern matching? We study this question through quantum circuit simulation as a training objective, where the model must predict the measurement probability distribution resulting from a sequence of quantum gate operations. We propose and compare two fine-tuning pipelines: (1) Supervised Fine-Tuning (SFT) on explicit gate-by-gate state-vector simulation traces, and (2) a two-stage SFT+Group Relative Policy Optimisation (GRPO) approach that sequentially applies SFT followed by GRPO with verifiable rewards. Our findings show that SFT achieves near-perfect in-distribution and gate-count extrapolation accuracy, significantly outperforming both the base model and the GPT-OSS-120B baseline. SFT+GRPO trades some in-distribution precision for better generalisation to larger qubit systems that SFT alone cannot handle. Both pipelines significantly outperform the baselines, demonstrating that targeted fine-tuning on explicit reasoning traces is an effective strategy for advancing quantum reasoning in LLMs.

量子推理大模型微调生成式AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。