arXiv:2501.06254cs.CLcs.AI2025-01ICLR被引 19

针对稀疏自编码器的语义评估,提出基于多义词的新评测方法

Rethinking Evaluation of Sparse Autoencoders through the Representation of Polysemous Words

论文配图:Rethinking Evaluation of Sparse Autoencoders through the Representation of Polysemous Words
图 1 · 摘自论文原文
  • 聚焦多义词,设计语义感知的稀疏自编码器评估体系
  • 发现传统最优指标组合可能混淆可解释性,无法有效提取单义特征
  • 揭示深层注意力模块对区分词语多义性的关键作用,适合模型可解释性研究者

稀疏自编码器(SAEs)作为提升大语言模型可解释性的有力工具,通过将多义神经元的复杂叠加映射为单义特征,构建稀疏词汇字典。然而,传统的均方误差与L0稀疏度指标忽略了对语义表征能力的评估——即是否能准确区分单词的不同含义。例如,学习到的稀疏特征能否识别同一词的多种语义尚不明确。本文提出一套以多义词为核心的SAE评估方法,分析单义特征的质量。研究发现,优化均方误差-稀疏度帕累托前沿的SAE可能反而混淆可解释性,未必能有效提取单义特征。对多义词的分析还揭示了大模型内部机制:深层网络与注意力模块在区分词义方面起关键作用。该语义导向的评估为理解多义性及现有SAE目标提供了新视角,有助于开发更实用的稀疏自编码器。

原文摘要 · Abstract (English)

Sparse autoencoders (SAEs) have gained a lot of attention as a promising tool to improve the interpretability of large language models (LLMs) by mapping the complex superposition of polysemantic neurons into monosemantic features and composing a sparse dictionary of words. However, traditional performance metrics like Mean Squared Error and L0 sparsity ignore the evaluation of the semantic representational power of SAEs -- whether they can acquire interpretable monosemantic features while preserving the semantic relationship of words. For instance, it is not obvious whether a learned sparse feature could distinguish different meanings in one word. In this paper, we propose a suite of evaluations for SAEs to analyze the quality of monosemantic features by focusing on polysemous words. Our findings reveal that SAEs developed to improve the MSE-L0 Pareto frontier may confuse interpretability, which does not necessarily enhance the extraction of monosemantic features. The analysis of SAEs with polysemous words can also figure out the internal mechanism of LLMs; deeper layers and the Attention module contribute to distinguishing polysemy in a word. Our semantics focused evaluation offers new insights into the polysemy and the existing SAE objective and contributes to the development of more practical SAEs.

稀疏自编码器可解释性多义词

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。