用量子编程竞赛数据评估大模型生成量子代码能力
QHackBench: Benchmarking Large Language Models for Quantum Code Generation Using PennyLane Hackathon Challenges
- 基于真实竞赛题构建量子代码评测集
- RAG增强模型在复杂算法上表现接近标准提示
- 多智能体迭代修正提升代码执行成功率
大型语言模型在代码生成方面展现出强大潜力,但在量子计算领域的应用仍待深入探索。本文基于量子编程黑客松(QHack)的真实挑战,构建了面向PennyLane的量子代码生成基准测试集QHackBench,评估模型在无提示与检索增强生成(RAG)两种策略下的表现。通过结构化评估框架,考察代码的功能正确性、语法有效性及运行成功度。结果显示,经增强的PennyLane数据集支持下,RAG模型在复杂量子算法任务中表现接近标准提示策略;同时引入多智能体迭代修正机制,进一步提升了代码执行成功率。为推动研究发展,本工作将公开发布QHackBench、评估框架与实验结果。
原文摘要 · Abstract (English)
Recent advances in Large Language Models (LLMs) have demonstrated strong potential in code generation, yet their effectiveness in quantum computing remains underexplored. This paper benchmarks LLMs for PennyLane-based quantum code generation using real-world challenges from the Quantum Hackathon (QHack). We introduce QHackBench, a novel benchmark dataset derived from QHack competitions, and evaluate model performance under vanilla prompting and Retrieval-Augmented Generation (RAG). Our structured evaluation framework assesses functional correctness, syntactic validity, and execution success across varying challenge difficulties. Results indicate that RAG-enhanced models, supplemented with an augmented PennyLane dataset, approximately generate similar results as the standard prompting, particularly in complex quantum algorithms. Additionally, we introduce a multi-agent evaluation pipeline that iteratively refines incorrect solutions, further enhancing execution success rates. To foster further research, we commit to publicly releasing QHackBench, along with our evaluation framework and experimental results, enabling continued advancements in AI-assisted quantum programming.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。