arXiv:2511.11711cs.LG2025-11被引 1

用敲扑法控制假阳性,让稀疏自编码器发现更可信的神经网络特征

Which Sparse Autoencoder Features Are Real? Model-X Knockoffs for False Discovery Rate Control

  • 引入模型- X 敲扑法,结合多重检验校正控制假发现率
  • 在 Pythia-70M 上选中129个特征,仅25%携带任务相关信号
  • 适合追求可复现、高可信度解释的机制可解释性研究者

尽管稀疏自编码器(SAE)对识别神经网络中的可解释特征至关重要,但仍难以区分真实计算模式与错误相关性。本文将 Model-X 敲扑法应用于 SAE 特征选择,利用 knock-off+ 方法在标准 Model-X 假设下(通过高斯近似潜在分布)实现有限样本下的假发现率(FDR)控制。通过对 Pythia-70M 进行情感分类分析,从512个高活跃度 SAE 潜在变量中筛选出129个特征,目标 FDR 为0.1。结果显示,被选中的特征中约25%包含任务相关信号,而75%不相关;所选特征的敲扑统计量比未选特征高出5.40倍。该方法结合 SAE 与多重检验感知推断,提供了一种可复现且有理论保障的可靠特征发现框架,推动了机制可解释性的基础发展。

原文摘要 · Abstract (English)

Although sparse autoencoders (SAEs) are crucial for identifying interpretable features in neural networks, it is still challenging to distinguish between real computational patterns and erroneous correlations. We introduce Model-X knockoffs to SAE feature selection, using knock-off+ to control the false discovery rate (FDR) with finite-sample guarantees under the standard Model-X assumptions (in our case, via a Gaussian surrogate for the latent distribution). We select 129 features at a target FDR q=0.1 after analyzing 512 high-activity SAE latents for sentiment classification using Pythia-70M. About 25% of the latents under examination carry task-relevant signal, whereas 75% do not, according to the chosen set, which displays a 5.40x separation in knockoff statistics compared to non-selected features. Our method offers a re-producible and principled framework for reliable feature discovery by combining SAEs with multiple-testing-aware inference, advancing the foundations of mechanistic interpretability.

稀疏自编码器可解释性假发现率机制分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。