arXiv:2502.09022cs.AI2025-02NAACL被引 16

通过自影响分析揭示Transformer模型的推理路径

Mechanistic Unveiling of Transformer Circuits: Self-Influence as a Key to Model Reasoning

  • 用电路分析和自影响函数追踪每个词元的重要性变化
  • 在IOI任务中发现模型遵循可解释的多步推理路径
  • 适合关注模型可解释性与推理机制的研究者

基于Transformer的语言模型取得了显著成功,但其内部机制因非线性交互和高维运算而高度不透明。尽管先前研究显示这些模型隐式嵌入了推理树,人类完成相同任务时会使用多种不同的逻辑推理方式,目前仍不清楚语言模型实际采用何种多步推理机制。本文旨在通过机制可解释性研究,探索语言模型在多步推理任务中的具体机制。我们采用电路分析与自影响函数,评估每个词元在推理过程中的重要性动态变化,从而映射出模型所采用的推理路径。该方法应用于GPT-2模型在预测任务(IOI)上的表现,结果表明底层电路揭示了人类可理解的推理过程。

原文摘要 · Abstract (English)

Transformer-based language models have achieved significant success; however, their internal mechanisms remain largely opaque due to the complexity of non-linear interactions and high-dimensional operations. While previous studies have demonstrated that these models implicitly embed reasoning trees, humans typically employ various distinct logical reasoning mechanisms to complete the same task. It is still unclear which multi-step reasoning mechanisms are used by language models to solve such tasks. In this paper, we aim to address this question by investigating the mechanistic interpretability of language models, particularly in the context of multi-step reasoning tasks. Specifically, we employ circuit analysis and self-influence functions to evaluate the changing importance of each token throughout the reasoning process, allowing us to map the reasoning paths adopted by the model. We apply this methodology to the GPT-2 model on a prediction task (IOI) and demonstrate that the underlying circuits reveal a human-interpretable reasoning process used by the model.

可解释性Transformer推理机制电路分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。