arXiv:2506.05239cs.LG2025-06被引 2

对比浅层与迭代稀疏自编码器,发现后者能更好提取相关特征。

Evaluating Sparse Autoencoders: From Shallow Design to Matching Pursuit

  • 用迭代匹配追踪改进浅层稀疏自编码器,提升特征提取能力。
  • 在MNIST上验证,新方法可逐步提升重构精度,且收敛稳定。
  • 适合研究神经网络可解释性、特征相关性建模的学者参考。

稀疏自编码器(SAEs)近年来成为神经表示可解释性的重要工具,基于字典学习原理从结构未知的神经表征中提取稀疏且可解释的特征。本文在控制环境下使用MNIST数据集评估SAEs,发现现有浅层架构隐含依赖准正交性假设,限制了对相关特征的提取能力。为突破此局限,我们比较了传统SAE与一种展开匹配追踪(MP-SAE)的迭代式SAE,该方法能通过残差引导提取层次化场景中出现的相关特征,并保证随着原子选择数量增加,重构误差单调下降。

原文摘要 · Abstract (English)

Sparse autoencoders (SAEs) have recently become central tools for interpretability, leveraging dictionary learning principles to extract sparse, interpretable features from neural representations whose underlying structure is typically unknown. This paper evaluates SAEs in a controlled setting using MNIST, which reveals that current shallow architectures implicitly rely on a quasi-orthogonality assumption that limits the ability to extract correlated features. To move beyond this, we compare them with an iterative SAE that unrolls Matching Pursuit (MP-SAE), enabling the residual-guided extraction of correlated features that arise in hierarchical settings such as handwritten digit generation while guaranteeing monotonic improvement of the reconstruction as more atoms are selected.

稀疏编码可解释性特征提取匹配追踪

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。