arXiv:2410.07961quant-phcs.DS2024-10NeurIPS被引 13

首个面向大模型设计量子算法的基准数据集,助力评估AI在量子编程中的能力。

QCircuitBench: A Large-Scale Dataset for Benchmarking Quantum Algorithm Design

  • 构建通用框架,将量子算法设计特征适配大语言模型
  • 涵盖25个算法、12万+数据点,覆盖基础到高级应用任务
  • 支持自动验证与交互式推理,适合量子计算与AI交叉研究者

量子计算因可实现远超经典计算的速度提升而备受关注,但其算法的设计与实现受制于量子力学的复杂性及对量子态的精确控制。尽管人工智能发展迅速,却缺乏专门用于此目的的数据集。本文提出QCircuitBench,首个用于评估大语言模型设计与实现量子算法能力的基准数据集。不同于传统代码生成,该任务具有高度灵活的设计空间。主要贡献包括:1)建立通用框架,将量子算法设计的关键特征适配大语言模型;2)实现从基础原语到高级应用的25个量子算法,覆盖3个任务套件,共120,290个数据点;3)内置自动验证与验证功能,支持无需人工干预的迭代评估与交互推理;4)初步微调实验显示其具备作为训练数据集的潜力。实验发现,大模型常呈现一致错误模式,且微调并不总优于少样本学习。整体而言,QCircuitBench为大模型驱动的量子算法设计提供了全面基准,并揭示了当前大模型在此领域的能力局限。

原文摘要 · Abstract (English)

Quantum computing is an emerging field recognized for the significant speedup it offers over classical computing through quantum algorithms. However, designing and implementing quantum algorithms pose challenges due to the complex nature of quantum mechanics and the necessity for precise control over quantum states. Despite the significant advancements in AI, there has been a lack of datasets specifically tailored for this purpose. In this work, we introduce QCircuitBench, the first benchmark dataset designed to evaluate AI's capability in designing and implementing quantum algorithms using quantum programming languages. Unlike using AI for writing traditional codes, this task is fundamentally more complicated due to highly flexible design space. Our key contributions include: 1. A general framework which formulates the key features of quantum algorithm design for Large Language Models. 2. Implementations for quantum algorithms from basic primitives to advanced applications, spanning 3 task suites, 25 algorithms, and 120,290 data points. 3. Automatic validation and verification functions, allowing for iterative evaluation and interactive reasoning without human inspection. 4. Promising potential as a training dataset through preliminary fine-tuning results. We observed several interesting experimental phenomena: LLMs tend to exhibit consistent error patterns, and fine-tuning does not always outperform few-shot learning. In all, QCircuitBench is a comprehensive benchmark for LLM-driven quantum algorithm design, and it reveals limitations of LLMs in this domain.

量子计算大模型算法设计基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。