arXiv:2607.01336quant-phcond-mat.dis-nn2026-07被引 1

用稀疏自编码器解析量子神经态的内部机制,发现其可被单特征精准调控物理量。

Mechanistic Interpretability and Causal Feature Steering of Neural Quantum States via Sparse Autoencoders

论文配图:Mechanistic Interpretability and Causal Feature Steering of Neural Quantum States via Sparse Autoencoders
图 1 · 摘自论文原文
  • 通过稀疏自编码器从残差流中提取无监督物理特征
  • 单特征干预可平滑调节可观测量,能量几乎不变
  • 为量子神经态提供可解释性诊断与可控干预工具

神经量子态(NQS)是一类极具表现力的变分波函数近似方法,但其内部机制尚不清晰:在仅基于变分目标训练的情况下,为何能准确捕捉未显式优化的物理可观测量?本文提出系统性方法,利用稀疏自编码器分析NQS的内部激活。从残差流中提取的特征与序参数、台阶磁化率、半链关联函数等物理可观测量强相关,覆盖基态表示与实时间演化。这些特征的发现完全无监督,未使用任何物理标签。进一步证明,对单个特征进行训练后干预,可平滑且单调地调控对应可观测量,同时保持变分能量基本不变。结果表明,NQS不仅是函数逼近器,更编码了丰富的可解释物理信息。本方法为NQS提供诊断与干预工具,奠定机械可解释性用于构建更可靠、透明的NQS的基础。

原文摘要 · Abstract (English)

Neural Quantum States (NQS) are a remarkably expressive class of variational ansätze for quantum many-body wavefunctions, yet little is understood about their internal mechanisms: trained on variational objectives alone, how do NQS accurately capture physical observables that they have never been explicitly optimized for? In this work, we present a systematic approach to analyze the internal activations of NQS using sparse autoencoders. We extract features from the residual stream and demonstrate that these features strongly correlate with physical observables such as order parameters, staggered magnetization, and half-chain correlators, across both ground state representation and real-time dynamics. Remarkably, the discovery of these features is entirely unsupervised, with no physical labels provided. We further establish that such features causally affect the corresponding observables predicted by NQS, by showing that targeted, post-training intervention on a \textit{single} feature smoothly and monotonically steers the corresponding observable, while leaving the variational energy nearly unchanged. These results demonstrate that NQS are not merely functional approximators, but encode rich, interpretable internal representations of physical information. Our approach provides both a diagnostic and an intervention tool for NQS, and serves as a foundation for using mechanistic interpretability towards more reliable, transparent NQS.

量子神经态可解释性自编码器因果干预

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。