证明量子电路在生成与预测任务上可超越经典大模型,且差距无法通过改进设计弥补。
Separating quantum circuits from classical LLMs

- 用常数深度量子电路构造不可被经典扩散模型近似采样的分布
- 发现需超线性宽度的变压器才能计算特定量子电路能处理的函数
- 揭示量子模型在语言任务中的本质优势,适合关注量子计算与大模型交叉的研究者
现代大型语言模型——包括变换器和扩散语言模型——围绕预测与生成两类核心任务构建。本文证明了低深度量子计算与相应资源受限的经典语言模型架构之间存在无条件分离。具体而言:1. 分布分离:我们构造了一个可由$ extsf{QNC}^0$电路(即由有界扇入门构成的常数深度量子电路族)采样的分布,任何具有浅层调度与去噪机制的常数轮扩散语言模型($ extsf{DLM}$)都无法在常数距离内采样该分布,即使允许子线性思维链及输出标记重写/遮蔽操作——这些正是现代$ extsf{DLM}$依赖的关键特性。2. 功能分离:我们展示了一个函数,该函数可在$ extsf{QNC}^0["log ext{log} n$]电路(即输入长度为$n$时深度为O$( ext{log log }n)$的$ extsf{QNC}^0$电路,后接一个经典$ extsf{AND}$门)上计算,但任何常数深度仅解码器的变换器若要计算该函数,其宽度必须达到 $n^{Ω(1)}$。我们的工作开启了大语言模型时代下量子优势研究的新方向。
原文摘要 · Abstract (English)
Modern large language models - transformers and diffusion language models - are built around two canonical algorithmic tasks: prediction and generation. We prove unconditional separations between low-depth quantum computation and the corresponding bounded-resource classical language-model architectures in both regimes. Concretely, we exhibit the following: 1. Distributional separation. We give a distribution that is sampleable by $\textsf{QNC}^0$ circuits (i.e., a family of constant-depth quantum circuits consisting of bounded fan-in gates) that no constant-round diffusion language model ($\textsf{DLM}$) with shallow scheduling and denoising can sample within constant distance, even when allowed sublinear chain-of-thought and output-token revision/remasking events, the very features modern $\textsf{DLM}$s rely on. 2. Functional separation. We exhibit a function computable in $\land \circ \textsf{QNC}^0[\log\log n]$ (i.e., a family of O$(\log\log n)$-depth $\textsf{QNC}^0$ circuits, where $n$ is the input length, followed by a single classical $\mathsf{AND}$ gate) such that any constant-depth decoder-only transformer computing the function must be large: it would have to have width $n^{Ω(1)}$. Together, our work initiates the study of quantum advantage in the era of large language models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。