arXiv:2603.26494quant-phcs.CL2026-03

量子语言模型靠纠缠记忆上下文,但噪声会破坏这种优势。

Entanglement as Memory: Mechanistic Interpretability of Quantum Language Models

  • 用纠缠追踪和门消融分析模型内部记忆机制。
  • 双量子比特模型通过纠缠编码上下文,三重因果检验显著支持此结论。
  • 真实硬件上纠缠策略因噪声退化,仅经典几何策略能存活。

量子语言模型在序列任务中表现优异,但其是否真正利用量子资源仍不明确。以往研究仅依赖终点指标,未探查模型内部学习的内存策略。本文首次对量子语言模型开展机制可解释性研究,结合因果门消融、纠缠追踪与密度矩阵干预,在长程依赖任务中验证。结果表明:单量子比特模型完全可经典模拟,收敛至与经典基线相同的几何策略;而含纠缠门的双量子比特模型则学习到表征不同的策略,将上下文编码于量子比特间纠缠(三组独立因果测试,p < 0.0001,d = 0.89)。在真实量子硬件上,仅经典几何策略幸存,纠缠策略受设备噪声影响退化至随机水平。这些发现揭示了噪声-表达力权衡机制,为量子语言模型的科学探索提供了新工具。

原文摘要 · Abstract (English)

Quantum language models have shown competitive performance on sequential tasks, yet whether trained quantum circuits exploit genuinely quantum resources -- or merely embed classical computation in quantum hardware -- remains unknown. Prior work has evaluated these models through endpoint metrics alone, without examining the memory strategies they actually learn internally. We introduce the first mechanistic interpretability study of quantum language models, combining causal gate ablation, entanglement tracking, and density-matrix interchange interventions on a controlled long-range dependency task. We find that single-qubit models are exactly classically simulable and converge to the same geometric strategy as matched classical baselines, while two-qubit models with entangling gates learn a representationally distinct strategy that encodes context in inter-qubit entanglement -- confirmed by three independent causal tests (p < 0.0001, d = 0.89). On real quantum hardware, only the classical geometric strategy survives device noise; the entanglement strategy degrades to chance. These findings open mechanistic interpretability as a tool for the science of quantum language models and reveal a noise-expressivity tradeoff governing which learned strategies survive deployment.

量子计算可解释性纠缠语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。