arXiv:2605.28567cs.LGcs.AI2026-05

用语义距离统一解决多层特征匹配与电路压缩难题

Semantic Optimal Transport for Sparse Autoencoder Feature Matching and Circuit Compression

论文配图:Semantic Optimal Transport for Sparse Autoencoder Feature Matching and Circuit Compression
图 1 · 摘自论文原文
  • 将特征表示为激活加权分布,而非单一解码向量
  • 通过沃尔什距离实现跨层特征精准匹配,准确率显著提升
  • 自动压缩大型特征电路为可解释超节点,适合模型可解释性研究

稀疏自编码器(SAEs)已成为解析语言模型的核心工具。然而,两个关键问题仍难以扩展:(1)跨多层匹配语义相似的特征;(2)将大型特征电路压缩为可解释的超节点。尽管这两类问题常被分别处理,我们发现它们均源于一个更根本的挑战——在不同激活流形上的SAE特征间估计语义距离。为此,我们提出一种分布式框架:每个特征不再由单一解码向量表示,而是由表达它的隐藏状态的激活加权分布来刻画。通过将这些分布投影至共享参考空间,并利用沃斯泰尔(Wasserstein)距离进行比较,该方法提供了一个统一的语义度量,用于跨层特征比对。理论证明该表示对激活缩放不变、扰动下稳定,并在有限样本条件下可恢复真实匹配。实验表明,该方法优于解码向量和基于大模型的基线方法,能捕捉相关特征间的细微功能差异。特别地,该方法可自动将大型特征电路压缩为可解释的超节点。

原文摘要 · Abstract (English)

Sparse autoencoders (SAEs) have become a central tool for interpreting language models. However, two key SAE analyses that remain difficult to scale are (1) matching semantically similar features across multi-layers and (2) compressing large feature circuits into interpretable supernodes. Although these have been treated as separate problems, we show that both are instances of a more fundamental challenge, which we frame as the estimation of semantic distances between SAE features that lie on different activation manifolds. We introduce a distributional framework for this problem, in which each feature is represented not by a single decoder vector like in the literature, but by an activation-weighted distribution over the hidden states that express it. By projecting these distributions into a shared reference space and comparing them with Wasserstein distance, our method provides a unified semantic metric for cross-layer feature comparison. We prove that our representation is invariant to activation rescaling, stable under perturbations, and recovers true matches under finite-sample margin conditions. Empirically, our method outperforms decoder-vector and LLM-based baselines and captures subtle functional distinctions between related features. Notably, our method compresses large feature circuits into interpretable supernodes automatically.

稀疏自编码器特征匹配可解释性分布建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。