arXiv:2608.14691cs.LGquant-ph2026-08

用复数相位态加速序列模型训练,效果比传统方法快一至三倍。

The Quantum Shortcut: Complex Phase-State Dynamics Reduce the Optimization Steps of Sequence Models

  • 用量子力学中的复数相位表示隐藏状态,替代传统实数表示
  • 在相同参数量下,训练步数减少至1/3到1/2,且长期性能更优
  • 适合追求高效训练的序列建模研究者,尤其关注优化效率

序列模型通常由骨干结构(如注意力或循环)区分,但本文关注的是其前置共性选择——表示基底:隐藏状态所依赖的数系及从状态映射到预测的函数形式。主流使用实数状态与仿射-软最大化读出;本文探索一种源自量子理论的复数替代方案,信息由状态相位携带,得分采用二次玻恩形式。先前工作证明该表示在理想条件下强于所有线性读出的实数模型;本文进一步检验其是否训练更快。放松阻碍部署的两个限制——精确幺正性和玻恩读出词汇,将该设计引入Mamba状态空间模型和基于注意力的Transformer。在253M参数、相同训练协议下,对三个字节级语料库进行实验,复数模型达到各验证损失所需优化步数约为实数模型的1/3(状态空间)和1/2(注意力)。此后两者分化:学习率预热结束后,状态空间模型优势持续扩大,在OpenWebText上从0.321降至0.354比特/字符,在FineWeb上从0.368降至0.396;而注意力模型的优势则逐渐衰减至零,表明其仅为早期训练效应。

原文摘要 · Abstract (English)

Sequence models are conventionally distinguished by their backbone, the mechanism that routes information across positions, such as attention or recurrence. This paper varies a choice that is prior to the backbone and shared by nearly all current models: the \emph{substrate}, the number system in which the hidden state is represented together with the form of the map from state to prediction. The prevailing substrate is a real-valued state with an affine--softmax readout; we study a complex-valued alternative drawn from the mathematics of quantum theory, in which information is carried by the phases of the state and scores are quadratic Born forms. Prior work proved an idealized version of this substrate representationally stronger than any real model with a linear readout; we ask whether it also trains faster. Relaxing the two properties that block deployment, exact unitarity and the Born vocabulary readout, we instantiate it in the Mamba state-space model and an attention-based Transformer. At 253M parameters, matched to within $0.02\%$ and trained under one fixed protocol on three byte-level corpora, the complex models reach every measured validation loss in approximately one third (state-space) and one half (attention) of the optimization steps of their real counterparts. The two backbones then diverge. Once the learning-rate warmup ends, the state-space advantage continues to widen, from $0.321$ to $0.354$ bits per character on OpenWebText and from $0.368$ to $0.396$ on FineWeb, which an artifact of the warmup ramp would not do; the attention advantage instead decays toward zero on every corpus, and is therefore an effect of early training.

序列建模复数表示优化加速量子启发

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。