用稀疏自编码器提取文本分类中可解释的关键概念,提升模型决策透明度。
Unveiling Decision-Making in LLMs for Text Classification : Extraction of influential and interpretable concepts with Sparse Autoencoders
- 设计专用分类头与稀疏损失,增强文本分类任务中的概念提取能力
- 在多个基准上验证,新方法显著提升概念的因果性与可解释性
- 适合关注大模型内部机制可解释性的研究者与开发者
稀疏自编码器(SAEs)已被成功用于探测大语言模型(LLMs)并从中提取可解释的概念,这些概念是神经元激活的线性组合,对应人类可理解的特征。本文研究了基于SAE的可解释性方法在句子分类领域的有效性,该领域此前相关探索较少。我们提出一种专为文本分类设计的新架构ClassifSAE,引入专用分类头并加入激活率稀疏性损失。在多个分类基准和骨干LLM上,将其与ConceptShap、独立成分分析、HI-Concept及标准TopK-SAE基线进行对比。此外,我们引入两种新指标,利用外部句子编码器评估概念解释的精度。实验结果表明,ClassifSAE在提取特征的因果性和可解释性方面均有显著提升。
原文摘要 · Abstract (English)
Sparse Autoencoders (SAEs) have been successfully used to probe Large Language Models (LLMs) and extract interpretable concepts from their internal representations. These concepts are linear combinations of neuron activations that correspond to human-interpretable features. In this paper, we investigate the effectiveness of SAE-based explainability approaches for sentence classification, a domain where such methods have not been extensively explored. We present a novel SAE-based model ClassifSAE tailored for text classification, leveraging a specialized classifier head and incorporating an activation rate sparsity loss. We benchmark this architecture against established methods such as ConceptShap, Independent Component Analysis, HI-Concept and a standard TopK-SAE baseline. Our evaluation covers several classification benchmarks and backbone LLMs. We further enrich our analysis with two novel metrics for measuring the precision of concept-based explanations, using an external sentence encoder. Our empirical results show that ClassifSAE improves both the causality and interpretability of the extracted features.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。