让稀疏自编码器特征更稳定,提升可解释性研究可靠性
Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs
- 用配对字典相关系数衡量特征一致性,指导模型设计
- 在大语言模型激活上实现0.80的高一致性水平
- 适合关注模型可解释性与结果复现的研究者
稀疏自编码器(SAEs)是机制可解释性中分解神经网络激活以识别可解释特征的重要工具。然而,不同训练运行间学到的特征不一致,挑战了找到标准特征集的目标,削弱了可解释性研究的可靠性和效率。本文主张,机制可解释性应优先考虑SAE中的特征一致性——即独立训练下收敛到相似特征集的能力。我们提出使用成对字典平均相关系数(PW-MCC)作为实际度量,并证明通过合适的架构选择可实现高水平一致性(在大语言模型激活上达0.80)。贡献包括阐明一致性优先的优势、通过模型生物体提供理论依据与合成验证,确认PW-MCC能可靠反映真实特征恢复;并扩展至真实大语言模型数据,发现高一致性与特征解释的语义相似性显著相关。呼吁社区系统性测量特征一致性,推动可解释性研究稳健积累。
原文摘要 · Abstract (English)
Sparse Autoencoders (SAEs) are a prominent tool in mechanistic interpretability (MI) for decomposing neural network activations into interpretable features. However, the aspiration to identify a canonical set of features is challenged by the observed inconsistency of learned SAE features across different training runs, undermining the reliability and efficiency of MI research. This position paper argues that mechanistic interpretability should prioritize feature consistency in SAEs -- the reliable convergence to equivalent feature sets across independent runs. We propose using the Pairwise Dictionary Mean Correlation Coefficient (PW-MCC) as a practical metric to operationalize consistency and demonstrate that high levels are achievable (0.80 for TopK SAEs on LLM activations) with appropriate architectural choices. Our contributions include detailing the benefits of prioritizing consistency; providing theoretical grounding and synthetic validation using a model organism, which verifies PW-MCC as a reliable proxy for ground-truth recovery; and extending these findings to real-world LLM data, where high feature consistency strongly correlates with the semantic similarity of learned feature explanations. We call for a community-wide shift towards systematically measuring feature consistency to foster robust cumulative progress in MI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。