arXiv:2411.13117cs.LG2024-11被引 19

改进稀疏自编码器推理,用更高效方法提升可解释性。

Compute Optimal Inference and Provable Amortisation Gap in Sparse Autoencoders

  • 分离编码与解码,用先进算法优化稀疏推理
  • 少量计算开销下实现显著更高的稀疏编码准确率
  • 适用于大模型激活分析,助力理解神经网络内部机制

近期研究显示,稀疏自编码器(SAEs)在揭示神经网络表征中的可解释特征方面具有潜力。然而,其简单的线性-非线性编码机制限制了稀疏推理的准确性。基于压缩感知理论,我们证明即使在可解情况下,SAE编码器也固有地无法实现准确稀疏推理。随后,我们解耦编码与解码过程,实证探索更复杂稀疏推理方法优于传统SAE编码器的条件。结果表明,在极小计算开销增加的情况下,稀疏代码的正确推理性能大幅提升。该方法在应用于大型语言模型的SAEs时同样有效,更具表达力的编码器能带来更高的可解释性。本工作为理解神经网络表征及分析大语言模型激活提供了新路径。

原文摘要 · Abstract (English)

A recent line of work has shown promise in using sparse autoencoders (SAEs) to uncover interpretable features in neural network representations. However, the simple linear-nonlinear encoding mechanism in SAEs limits their ability to perform accurate sparse inference. Using compressed sensing theory, we prove that an SAE encoder is inherently insufficient for accurate sparse inference, even in solvable cases. We then decouple encoding and decoding processes to empirically explore conditions where more sophisticated sparse inference methods outperform traditional SAE encoders. Our results reveal substantial performance gains with minimal compute increases in correct inference of sparse codes. We demonstrate this generalises to SAEs applied to large language models, where more expressive encoders achieve greater interpretability. This work opens new avenues for understanding neural network representations and analysing large language model activations.

稀疏编码可解释性大模型分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。