arXiv:2506.03093cs.LG2025-06NeurIPS被引 45

提出新型稀疏自编码器MP-SAE,捕捉神经网络中层级化、非线性的抽象特征。

From Flat to Hierarchical: Extracting Sparse Representations with Matching Pursuit

  • 将匹配追踪算法重构为残差引导的逐步编码,实现层级特征建模
  • 在合成与真实数据上验证了现有方法无法准确捕捉条件正交特征
  • 可自适应稀疏性,适合多模态模型中深层结构的可解释性研究

受神经网络表征以线性可访问、近似正交方向编码抽象可解释特征这一假设驱动,稀疏自编码器(SAEs)已成为可解释性研究的热门工具。然而,近期工作揭示了模型表征中存在超出该假设的现象,表现为层级化、非线性和多维特征。这引发疑问:现有SAE是否无法反映与其假设相悖的特征?若否,避免这种不匹配能否帮助识别这些特征并深化对神经网络表征的理解?为此,本文采用构建式方法,将流行的匹配追踪(MP)算法从稀疏编码视角重构为设计MP-SAE——一种将编码器展开为一系列残差引导步骤的SAE,使其能捕捉层级化与非线性可访问特征。在合成与自然数据混合设置下对比现有SAE,我们发现:(i) 层级概念诱导出条件正交特征,而现有SAE无法忠实捕捉;(ii) MP-SAE的非线性编码步骤恢复出高度有意义的特征,帮助我们揭示视觉-语言模型中不同模态表征空间看似对立却共享的结构,从而证明‘有用特征仅线性可访问’这一假设不足。此外,MP-SAE的序列编码机制还带来推理时自适应稀疏性的额外优势。总体而言,我们认为结果支持可解释性应始于表征现象学,方法需基于契合其本质的假设。

原文摘要 · Abstract (English)

Motivated by the hypothesis that neural network representations encode abstract, interpretable features as linearly accessible, approximately orthogonal directions, sparse autoencoders (SAEs) have become a popular tool in interpretability. However, recent work has demonstrated phenomenology of model representations that lies outside the scope of this hypothesis, showing signatures of hierarchical, nonlinear, and multi-dimensional features. This raises the question: do SAEs represent features that possess structure at odds with their motivating hypothesis? If not, does avoiding this mismatch help identify said features and gain further insights into neural network representations? To answer these questions, we take a construction-based approach and re-contextualize the popular matching pursuits (MP) algorithm from sparse coding to design MP-SAE -- an SAE that unrolls its encoder into a sequence of residual-guided steps, allowing it to capture hierarchical and nonlinearly accessible features. Comparing this architecture with existing SAEs on a mixture of synthetic and natural data settings, we show: (i) hierarchical concepts induce conditionally orthogonal features, which existing SAEs are unable to faithfully capture, and (ii) the nonlinear encoding step of MP-SAE recovers highly meaningful features, helping us unravel shared structure in the seemingly dichotomous representation spaces of different modalities in a vision-language model, hence demonstrating the assumption that useful features are solely linearly accessible is insufficient. We also show that the sequential encoder principle of MP-SAE affords an additional benefit of adaptive sparsity at inference time, which may be of independent interest. Overall, we argue our results provide credence to the idea that interpretability should begin with the phenomenology of representations, with methods emerging from assumptions that fit it.

稀疏编码可解释性层级特征匹配追踪

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。