arXiv:2506.04859cs.LGcs.AI2025-06ICML被引 2

提出新型混合自编码器,更好捕捉数据低维结构并生成更稀疏表示。

Sparse Autoencoders, Again?

  • 融合确定性与随机编码机制,突破传统稀疏自编码器局限。
  • 在真实与合成数据上实现更低的潜在维度估计误差和更稀疏表示。
  • 适用于语言模型激活模式与图像数据,优于同等容量的SAE、VAE及扩散模型。

稀疏自编码器(SAEs)虽能建模数据中的低维潜在结构,如大语言模型激活的相关模式或自然图像流形,但其核心架构近几十年变化甚微,仍依赖经典深度编解码结构与确定性稀疏正则化。变分自编码器(VAE)虽引入随机编码以生成稀疏表示,但同样存在未被充分重视的缺陷。本文揭示了典型SAE与类似任务下的VAE的理论局限,并提出一种混合模型以克服这些瓶颈。理论上,该模型全局最优解可恢复分布在多个流形并集上的结构化数据;实验上,在合成与真实数据集上验证了其在准确估计流形维度、生成更稀疏潜在表示方面的能力,同时保持重建误差不降。总体而言,本方法在图像与语言模型激活模式等场景中,性能超越同等容量的SAE、VAE及部分扩散模型。

原文摘要 · Abstract (English)

Is there really much more to say about sparse autoencoders (SAEs)? Autoencoders in general, and SAEs in particular, represent deep architectures that are capable of modeling low-dimensional latent structure in data. Such structure could reflect, among other things, correlation patterns in large language model activations, or complex natural image manifolds. And yet despite the wide-ranging applicability, there have been relatively few changes to SAEs beyond the original recipe from decades ago, namely, standard deep encoder/decoder layers trained with a classical/deterministic sparse regularizer applied within the latent space. One possible exception is the variational autoencoder (VAE), which adopts a stochastic encoder module capable of producing sparse representations when applied to manifold data. In this work we formalize underappreciated weaknesses with both canonical SAEs, as well as analogous VAEs applied to similar tasks, and propose a hybrid alternative model that circumvents these prior limitations. In terms of theoretical support, we prove that global minima of our proposed model recover certain forms of structured data spread across a union of manifolds. Meanwhile, empirical evaluations on synthetic and real-world datasets substantiate the efficacy of our approach in accurately estimating underlying manifold dimensions and producing sparser latent representations without compromising reconstruction error. In general, we are able to exceed the performance of equivalent-capacity SAEs and VAEs, as well as recent diffusion models where applicable, within domains such as images and language model activation patterns.

自编码器稀疏表示流形学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。