arXiv:2602.06852quant-phcs.AI2026-02

用量子方法追踪大模型内部记忆回溯机制,发现不同模型有截然不同的工作原理。

The Quantum Sieve Tracer: A Hybrid Framework for Layer-Wise Activation Tracing in Large Language Models

  • 结合经典与量子方法,分两阶段定位关键层并映射注意力激活
  • 发现Llama-3.2的第9层抑制干扰反而提升记忆准确率,反直觉但有效
  • 适合研究模型可解释性、注意力机制和量子计算应用的科研人员

机制可解释性旨在逆向解析大语言模型的内部计算过程,但分离稀疏语义信号与高维多义噪声仍是重大挑战。本文提出量子筛子追踪器(Quantum Sieve Tracer),一种混合量子-经典框架,用于刻画事实回忆回路。我们构建模块化流程:先用经典因果追踪定位关键层,再将特定注意力头激活映射至指数级庞大的量子希尔伯特空间。基于开源模型Meta Llama-3.2-1B与Alibaba Qwen2.5-1.5B-Instruct进行两阶段分析,揭示基础架构差异:Qwen第7层为典型回忆枢纽,而Llama第9层表现为干扰抑制电路——删除该层指定注意力头反而提升事实回忆准确率。结果表明,量子核函数能有效区分建设性(回忆)与还原性(抑制)机制,提供高分辨率工具以解析注意力的细粒度拓扑结构。

原文摘要 · Abstract (English)

Mechanistic interpretability aims to reverse-engineer the internal computations of Large Language Models (LLMs), yet separating sparse semantic signals from high-dimensional polysemantic noise remains a significant challenge. This paper introduces the Quantum Sieve Tracer, a hybrid quantum-classical framework designed to characterize factual recall circuits. We implement a modular pipeline that first localizes critical layers using classical causal tracing, then maps specific attention head activations into an exponentially large quantum Hilbert space. Using open-weight models (Meta Llama-3.2-1B and Alibaba Qwen2.5-1.5B-Instruct), we perform a two-stage analysis that reveals a fundamental architectural divergence. While Qwen's layer 7 circuit functions as a classic Recall Hub, we discover that Llama's layer 9 acts as an Interference Suppression circuit, where ablating the identified heads paradoxically improves factual recall. Our results demonstrate that quantum kernels can distinguish between these constructive (recall) and reductive (suppression) mechanisms, offering a high-resolution tool for analyzing the fine-grained topology of attention.

可解释性注意力机制量子计算大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。