arXiv:2505.08080cs.LGcs.AI2025-05EMNLP被引 4

用输出梯度识别对大模型影响最大的隐藏特征。

Beyond Input Activations: Identifying Influential Latents by Gradient Sparse Autoencoders

  • 结合输出梯度筛选关键隐藏变量
  • 发现激活不等于重要,仅高因果影响力特征有效
  • 适合想精准控制LLM行为的研究者

稀疏自编码器(SAEs)已成为解释和操控大语言模型(LLMs)内部表征的强大工具。然而,现有分析方法通常仅依赖输入侧激活,未考虑每个隐藏特征对模型输出的因果影响。本文基于两个核心假设:(1) 激活的隐藏特征对模型输出的贡献并不均等;(2) 仅具有高因果影响力的隐藏特征才适用于模型操控。为此,我们提出梯度稀疏自编码器(GradSAE),一种简单有效的方法,通过引入输出侧梯度信息,识别最具影响力的隐藏特征。

原文摘要 · Abstract (English)

Sparse Autoencoders (SAEs) have recently emerged as powerful tools for interpreting and steering the internal representations of large language models (LLMs). However, conventional approaches to analyzing SAEs typically rely solely on input-side activations, without considering the causal influence between each latent feature and the model's output. This work is built on two key hypotheses: (1) activated latents do not contribute equally to the construction of the model's output, and (2) only latents with high causal influence are effective for model steering. To validate these hypotheses, we propose Gradient Sparse Autoencoder (GradSAE), a simple yet effective method that identifies the most influential latents by incorporating output-side gradient information.

自编码器模型可解释性大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。