arXiv:2605.12874cs.LG2026-05被引 2

发现自编码器解释存在'描述碰撞',多个特征共用同一解释,影响可解释性评估。

Descriptive Collision in Sparse Autoencoder Auto-Interpretability: When One Explanation Describes Many Features

  • 识别出多个SAE特征共享同一自然语言解释的现象
  • 82.1%的特征与至少一个其他特征共享解释,最常见解释覆盖101个特征
  • 提出修正指标,惩罚无法区分特征的模糊解释,提升评估可靠性

稀疏自编码器(SAE)已成为分解语言模型激活以获得可解释特征的标准工具,自动化可解释性流程通常为每个特征分配简短的自然语言解释。现有批评多聚焦于一词多义(单特征多重含义)或解释对激活的预测能力。本文识别出一种结构上不同的新问题:描述碰撞——多个不同特征采用相同解释。重新分析最大公开的真人标注SAE特征数据集(Marks et al., 2025),涵盖Gemma 2 2B和Pythia 70M模型共722个标注特征,发现平均解释字符串被3.07个特征重复使用;82.1%的特征至少与一个其他特征共享解释;最常见解释“plural nouns”覆盖101个不同特征,跨越18层和4个模型组件。信息论分析显示,平均解释仅能分辨70%的特征身份。我们形式化了‘判别性’属性,证明当前检测式自动可解释性评分对碰撞不敏感,并提出两种互补修正指标——碰撞调整后的检测评分与判别性评分——明确惩罚无法区分特征邻域的解释。该问题独立于且叠加于此前已知的自动可解释性失败模式,忽略它会使报告的可解释性高估量达到标识特征所需比特数的大约三分之一。

原文摘要 · Abstract (English)

Sparse autoencoders (SAEs) are now standard tools for decomposing language model activations into interpretable features, and automated interpretability pipelines routinely assign each feature a short natural-language explanation. Existing critiques of this practice focus on polysemanticity -- one feature with many meanings -- or on whether explanations predict activations. We identify a complementary, structurally distinct problem we call descriptive collision: many distinct SAE features admit the same explanation. Reanalyzing the largest publicly-available dataset of human-annotated SAE features (Marks et al., 2025), comprising 722 annotated features across Gemma 2 2B and Pythia 70M, we find that the mean annotation string is reused across 3.07 features; 82.1% of features share their annotation with at least one other feature; and the single most common annotation string ("plural nouns") labels 101 distinct features spanning 18 layers and four model components. Information-theoretically, the average annotation resolves only 70% of feature identity. We formalize a property called discrimination, prove that current detection-style auto-interpretability scoring is invariant to collision, and propose two complementary corrective metrics -- collision-adjusted detection and discrimination scoring -- that explicitly penalize explanations that fail to distinguish a feature from its neighbors. The collision problem is independent of, and additive with, previously identified failure modes of auto-interpretability; ignoring it inflates reported feature interpretability by a quantity equal to roughly one-third of the bits required to identify a feature.

可解释性自编码器特征解析语义重叠

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。