arXiv:2503.05613cs.LGcs.AI2025-03EMNLP综述被引 78

用稀疏自编码器解析大模型内部机制,让黑箱变透明

A Survey on Sparse Autoencoders: Interpreting the Internal Mechanisms of Large Language Models

  • 通过稀疏自编码器分解大模型中的复杂特征
  • 揭示特征与语言行为间的可解释关联
  • 适合研究大模型可解释性的学者和工程师

大语言模型(LLMs)已深刻改变自然语言处理领域,但其内部工作机制仍高度不透明。近年来,机制可解释性成为研究热点,旨在揭示大模型的内在运作原理。在多种机制可解释性方法中,稀疏自编码器(SAEs)因其能将大模型中复杂的叠加特征解耦为更可解释的成分而展现出巨大潜力。本文系统综述了用于理解大模型内部机制的稀疏自编码器技术。主要贡献包括:(1) 探讨SAE的技术框架,涵盖基础架构、设计改进及高效训练策略;(2) 分析不同特征解释方法,分为基于输入和基于输出两类;(3) 讨论评估SAE性能的方法,涵盖结构与功能指标;(4) 探索SAE在理解与操控大模型行为中的实际应用。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have transformed natural language processing, yet their internal mechanisms remain largely opaque. Recently, mechanistic interpretability has attracted significant attention from the research community as a means to understand the inner workings of LLMs. Among various mechanistic interpretability approaches, Sparse Autoencoders (SAEs) have emerged as a promising method due to their ability to disentangle the complex, superimposed features within LLMs into more interpretable components. This paper presents a comprehensive survey of SAEs for interpreting and understanding the internal workings of LLMs. Our major contributions include: (1) exploring the technical framework of SAEs, covering basic architecture, design improvements, and effective training strategies; (2) examining different approaches to explaining SAE features, categorized into input-based and output-based explanation methods; (3) discussing evaluation methods for assessing SAE performance, covering both structural and functional metrics; and (4) investigating real-world applications of SAEs in understanding and manipulating LLM behaviors.

可解释性大模型稀疏编码

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。