用稀疏自编码器揭示物理神经网络的隐藏表征,发现其编码了可因果验证的物理特征。
PhysSAE: Mechanistic Interpretability with Sparse Autoencoders

- 在PINN的中间层训练稀疏自编码器,通过直接干预检测特征的因果作用。
- 发现的特征与物理可观测量高度相关,且因果影响更集中(1.2–4.2倍于PCA/ICA)。
- 适合关注科学机器学习可解释性的研究者,尤其关心物理规律如何被模型隐式编码。
物理信息神经网络(PINNs)将偏微分方程残差嵌入训练过程,但其内部表征仍不透明:无法确定隐藏层编码了哪些物理特征,或这些特征是否具有局部因果作用。我们提出PhysSAE,一种机制可解释性框架,对六类偏微分方程、每类3个PINN种子和3个SAE种子的样本,在PINN的倒数第二层激活上训练过完备稀疏自编码器(SAEs),并通过直接因果干预原冻结隐藏状态 $h_{ ext{cf}} = h - \alpha z_k d_k$ 来评估字典原子,完全跳过SAE解码器。结果表明:(i) 发现的SAE原子与独立定义的物理可观测量高度一致(最大皮尔逊相关系数 $|r|=0.951$,显著高于置换零模型);(ii) 高度对齐原子的消融因果足迹比标准方法(PCA/ICA)的空间集中度高1.2–4.2倍;(iii) 高度对齐原子在结构化物理概念上的因果定位性能优于随机对照组(ESF$_{80}$优势0.04–0.44)。双原子双边表示使概念回归 $R^2$ 提升 $ riangle R^2=0.05 ext{-}0.15$,而随机对则下降最多0.60。结果表明,PINNs具备稀疏且具物理结构的潜在表征,可事后识别并因果探查,为可解释性驱动的科学机器学习开辟路径。
原文摘要 · Abstract (English)
Physics-Informed Neural Networks (PINNs) embed PDE residuals into neural network training, but their internal representations remain opaque: it is unknown what physical features their hidden layers encode or whether those features have a localized causal role. We present PhysSAE, a mechanistic interpretability framework that trains overcomplete sparse autoencoders (SAEs) on PINN penultimate-layer activations and evaluates dictionary atoms through direct causal intervention in the original frozen hidden state: $h_{\mathrm{cf}} = h - \alpha z_k d_k$, bypassing the SAE decoder entirely. Across six PDE families, with 3 PINN seeds and 3 SAE seeds each---we show that (i) Our discovered SAE atoms align with independently-defined physical observables (max Pearson $|r|=0.951$, always $\gg$ permutation null), (ii) the causal footprint of top-aligned atom ablation is 1.2--4.2$\times$ more spatially concentrated canonical than PCA or ICA interventions, and (iii) top-aligned atoms outperform matched random controls on causal localization for structured physical concepts (ESF$_{80}$ advantage 0.04-0.44). Two-atom bilateral representations improve concept regression R$^2$ by $\Delta R^2\!=\!0.05\text{-}0.15$ over single atoms, while random pairs decrease it by up to 0.60. These results demonstrate that PINNs develop sparse, physically structured latent representations that can be identified and causally interrogated post-hoc, opening a path toward interpretability-aware scientific machine learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。