提出无标签度量方法,精准评估稀疏自编码器的语义单一性。
Measuring Monosemanticity in Sparse Autoencoders via Latent Activation Coherence

- 基于二值化潜变量激活的一致性定义语义单一性,无需外部标签或嵌入模型。
- 在多种视觉与视觉-语言模型上验证,该方法受编码器几何影响更小。
- 适用于研究SAS训练动态,且对概念删除效果有更强预测力。
在可解释人工智能中,机制可解释性利用稀疏自编码器(SAE)从神经表示中提取更易理解的特征。然而,评估其语义单一性——即解释质量——仍具挑战。现有度量依赖外部概念标签或预训练嵌入模型,易受编码器几何影响。本文提出无需标签的Tversky语义单一性评分(TMS),将语义单一性操作化为二值化SAE潜变量激活集的一致性,不依赖外部嵌入编码器。我们在DINOv3、CLIP、BLIP2等预训练视觉与视觉-语言模型的特征上,针对两种常见SAE训练范式(TopK、BatchTopK)、多个稀疏度水平与扩展因子,评估了TMS。结果表明,相比基于嵌入的替代指标,TMS受编码器各向异性影响更小,同时与既有语义单一性指标保持一致。TMS还揭示了不同基础模型间的SAE训练动态差异。此外,在编码器各向异性条件下,TMS对探测器驱动的概念删除有效性具有更强指示作用,而在其他情况下表现相当。
原文摘要 · Abstract (English)
Within Explainable Artificial Intelligence, mechanistic interpretability uses Sparse Autoencoders (SAEs) to extract more interpretable features from neural representations. However, assessing their monosemanticity, and thus explanation quality, remains challenging. Existing metrics require external concept labels or depend on pretrained embedding models, making them sensitive to encoder's geometry. We introduce the Tversky Monosemanticity Score (TMS), a label-free metric that operationalizes monosemanticity as activation-set coherence of binarized SAE latents, and does not require external embedding encoders. We evaluate TMS on SAEs trained on features from pretrained vision and vision-language models (DINOv3, CLIP, BLIP2), two common SAE regimes (TopK, BatchTopK), multiple sparsity levels, and expansion factors. Our results show that TMS is less affected by encoder anisotropy than its embedding-based alternative, while remaining aligned with established monosemanticity indicators. TMS also reveals distinct SAE training dynamics across base models. Moreover, under encoder anisotropy, TMS provides a stronger indication of probe-based concept deletion effectiveness, while being competitive otherwise.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。