arXiv:2509.03738cs.LGcs.AI2025-09被引 1

用函数空间稀疏编码,让模型不仅能识别概念,还能理解其在输入中的表达方式。

Mechanistic Interpretability with Sparse Autoencoder Neural Operators

  • 将概念表示为函数而非数值,实现对概念表达位置和方式的建模
  • 在视觉数据上学习到局部模式,且在不同稀疏度下保持稳定
  • 可处理训练外分辨率,适合有空间结构的数据

我们提出稀疏自编码器神经算子(SAE-NOs),一种在函数空间中操作的新型稀疏自编码器,而非固定维度的欧几里得表示。我们形式化了函数表示假说:数据可通过结构化函数的稀疏组合来解释。与传统 SAEs 用标量激活表示概念不同,SAE-NOs 将概念参数化为函数,能捕捉概念的存在性及其在输入域中的表达方式和位置。这通过联合稀疏性实现:概念稀疏性选择活跃概念,域稀疏性控制其表达位置。我们使用傅里叶神经算子实例化该框架(SAE-FNOs),在傅里叶域中将概念参数化为积分算子。这种函数与频域参数化在数据具有跨尺度空间结构或概念为频率结构时尤为有利。我们在视觉数据上评估 SAE-FNO,结果表明其学习局部模式,更高效地使用概念,并在不同稀疏度下表现出稳定的概念特性。我们还证明,当输入域大小变化时,SAE-FNO 能适应并泛化至训练中未见的离散化分辨率,而标准 SAEs 在此失效。此外,我们引入向 SAE 的提升机制,理论与实证均表明其作为预条件器可加速优化。总体而言,从向量值到函数参数化,结合概念与域稀疏性,使 SAEs 不仅能表示概念存在,更能建模结构化的概念表达,凸显参数化的重要性。

原文摘要 · Abstract (English)

We introduce sparse autoencoder neural operators (SAE-NOs), a new class of sparse autoencoders that operate in function spaces rather than fixed-dimensional Euclidean representations. We formalize the functional representation hypothesis, where data are explained through sparse compositions of structured functions. Unlike standard SAEs that represent concepts with scalar activations, SAE-NOs parameterize concepts as functions, enabling representations that capture not only a concept's presence, but also how and where it is expressed across the input domain. We achieve this through joint sparsity: concept sparsity selects active concepts, while domain sparsity governs where they are expressed. We instantiate this framework using Fourier neural operators (SAE-FNOs), parameterizing concepts as integral operators in the Fourier domain. This functional and spectral parameterization is particularly advantageous when data exhibit spatial structure across scales or when concepts are frequency-structured. We characterize SAE-FNO on vision data and demonstrate that it learns localized patterns, uses concepts more efficiently, and exhibits stable concept characteristics across sparsity levels. We further show that SAE-FNO adapts to changes in domain size and generalizes across discretizations, operating at resolutions beyond those seen during training, where standard SAEs fail. We also introduce lifting into SAEs and show theoretically and empirically that it acts as a preconditioner that accelerates optimization. Overall, our results show that moving from vector-valued to functional parameterizations, with concept and domain sparsity, extends SAEs from representing concept presence to modeling structured concept expression, highlighting the importance of parameterization.

可解释性稀疏编码函数空间神经算子

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。