arXiv:2511.13981cs.LGcs.AI2025-11中稿 · AAAI被引 2

数据去相关可显著提升稀疏自编码器的可解释性

Data Whitening Improves Sparse Autoencoder Learning

  • 对输入激活做PCA去相关预处理,改善优化路径
  • 在多个模型和稀疏度下,可解释性指标普遍提升
  • 适合追求特征可解释性的研究者使用

稀疏自编码器(SAEs)已成为从神经网络激活中学习可解释特征的有前景方法。然而,由于输入数据存在相关性,SAE训练的优化景观可能具有挑战性。我们证明,在输入激活上应用PCA去相关(一种经典稀疏编码中的标准预处理技术)能提升SAE在多个指标上的表现。通过理论分析和仿真,我们发现去相关使优化景观更凸,更易导航。我们在多种模型架构、宽度和稀疏度下评估了ReLU和Top-K SAEs。在SAEBench这一全面的SAE基准测试中,去相关一致性地提升了可解释性指标,包括稀疏探测准确率和特征解耦度,尽管重建质量略有下降。结果挑战了可解释性与最优稀疏-保真权衡一致的假设,表明当可解释性优先于完美重建时,应将去相关作为默认预处理步骤。

原文摘要 · Abstract (English)

Sparse autoencoders (SAEs) have emerged as a promising approach for learning interpretable features from neural network activations. However, the optimization landscape for SAE training can be challenging due to correlations in the input data. We demonstrate that applying PCA Whitening to input activations -- a standard preprocessing technique in classical sparse coding -- improves SAE performance across multiple metrics. Through theoretical analysis and simulation, we show that whitening transforms the optimization landscape, making it more convex and easier to navigate. We evaluate both ReLU and Top-K SAEs across diverse model architectures, widths, and sparsity regimes. Empirical evaluation on SAEBench, a comprehensive benchmark for sparse autoencoders, reveals that whitening consistently improves interpretability metrics, including sparse probing accuracy and feature disentanglement, despite minor drops in reconstruction quality. Our results challenge the assumption that interpretability aligns with an optimal sparsity--fidelity trade-off and suggest that whitening should be considered as a default preprocessing step for SAE training, particularly when interpretability is prioritized over perfect reconstruction.

稀疏编码可解释性数据预处理特征解耦

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。