arXiv:2508.20907quant-phcs.AI2025-08被引 9

用量子验证奖励训练模型,提升量子代码生成质量

Quantum Verifiable Rewards for Post-Training Qiskit Code Assistant

  • 通过量子硬件验证生成奖励信号,指导模型优化代码
  • 在Qiskit-HumanEval-hard上超越最强开源基线
  • 结合数据合成与强化学习,适合量子编程研究者

Qiskit 是一个开源量子计算框架,支持用户设计、模拟并在真实量子硬件上运行量子电路。本文探索了后训练技术用于大模型辅助编写 Qiskit 代码。提出利用量子验证作为确保代码质量和可执行性的有效方法。为此,我们构建了一个合成数据流水线,生成量子问题与单元测试对,并以此创建偏好数据,用于 DPO 对齐。同时采用 GRPO 训练模型,利用量子硬件提供的可验证奖励信号。最佳模型结合 DPO 与 GRPO,在具有挑战性的 Qiskit-HumanEval-hard 基准上超越现有最强开源基线。

原文摘要 · Abstract (English)

Qiskit is an open-source quantum computing framework that allows users to design, simulate, and run quantum circuits on real quantum hardware. We explore post-training techniques for LLMs to assist in writing Qiskit code. We introduce quantum verification as an effective method for ensuring code quality and executability on quantum hardware. To support this, we developed a synthetic data pipeline that generates quantum problem-unit test pairs and used it to create preference data for aligning LLMs with DPO. Additionally, we trained models using GRPO, leveraging quantum-verifiable rewards provided by the quantum hardware. Our best-performing model, combining DPO and GRPO, surpasses the strongest open-source baselines on the challenging Qiskit-HumanEval-hard benchmark.

量子编程LLM强化学习代码生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。