提出一种无需模型结构信息的词元重要性评估方法
NormXLogit: The Head-on-Top Never Lies
- 基于输入输出表征计算词元重要性,不依赖具体模型结构
- 在忠实性上优于梯度法,在层间解释上媲美专用方法
- 适合需要跨架构可解释性的研究者快速分析模型决策
随着大型语言模型不断涌现,考虑具有模型无关性的可解释性方法的价值愈发重要。尽管现有大模型可解释性研究取得进展,但多数方法依赖复杂且特定于模型的机制,计算成本高。为此,我们提出NormXLogit,一种用于评估单个输入词元重要性的新方法。该方法基于每个词元的输入和输出表征。首先,我们证明在大模型预训练过程中,词嵌入的范数能有效捕捉词元重要性。其次,揭示了词元重要性与其表征与模型最终预测的相似程度之间存在显著关系。大量分析表明,该方法在忠实性方面优于现有梯度法,并在层间解释性能上与领先的架构特定技术相当。
原文摘要 · Abstract (English)
With new large language models (LLMs) emerging frequently, it is important to consider the potential value of model-agnostic approaches that can provide interpretability across a variety of architectures. While recent advances in LLM interpretability show promise, many rely on complex, model-specific methods with high computational costs. To address these limitations, we propose NormXLogit, a novel technique for assessing the significance of individual input tokens. This method operates based on the input and output representations associated with each token. First, we demonstrate that during the pre-training of LLMs, the norms of word embeddings effectively capture token importance. Second, we reveal a significant relationship between a token's importance and the extent to which its representation can resemble the model's final prediction. Extensive analyses reveal that our approach outperforms existing gradient-based methods in terms of faithfulness and offers competitive performance in layer-wise explanations compared to leading architecture-specific techniques.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。