arXiv:2411.18895cs.LGcs.CL2024-11被引 14

用自动化任务评估稀疏自编码器的可解释性效果。

Evaluating Sparse Autoencoders on Targeted Concept Erasure Tasks

  • 用大模型替代人工标注,自动判断特征是否无关并移除干扰线索。
  • 提出新指标TPP,能衡量模型对相似概念的解耦能力。
  • 适用于对比不同训练参数和架构的稀疏自编码器性能。

稀疏自编码器(SAEs)是一种将神经网络激活分解为可解释单元的可解释性技术。然而,其发展长期受限于高质量评估指标的缺乏,以往工作多依赖无监督代理指标。本文基于Marks等人(2024,《稀疏特征电路》)提出的下游任务SHIFT,引入一系列评估方法:在SHIFT中,通过擦除人类标注为任务无关的SAE特征来消除分类器中的虚假线索。我们将其自动化,用大语言模型(LLM)替代人工标注。此外,提出目标探测扰动(TPP)指标,量化SAE分离相似概念的能力,使SHIFT可扩展至更广泛的数据集。我们在多个开源模型上应用SHIFT与TPP,结果表明这些指标能有效区分不同超参数与架构的SAE性能。

原文摘要 · Abstract (English)

Sparse Autoencoders (SAEs) are an interpretability technique aimed at decomposing neural network activations into interpretable units. However, a major bottleneck for SAE development has been the lack of high-quality performance metrics, with prior work largely relying on unsupervised proxies. In this work, we introduce a family of evaluations based on SHIFT, a downstream task from Marks et al. (Sparse Feature Circuits, 2024) in which spurious cues are removed from a classifier by ablating SAE features judged to be task-irrelevant by a human annotator. We adapt SHIFT into an automated metric of SAE quality; this involves replacing the human annotator with an LLM. Additionally, we introduce the Targeted Probe Perturbation (TPP) metric that quantifies an SAE's ability to disentangle similar concepts, effectively scaling SHIFT to a wider range of datasets. We apply both SHIFT and TPP to multiple open-source models, demonstrating that these metrics effectively differentiate between various SAE training hyperparameters and architectures.

可解释性稀疏编码评估指标大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。