arXiv:2502.11356cs.LGcs.AI2025-02被引 30

用稀疏自编码器解析大模型如何理解指令并实现精准控制

SAIF: A Sparse Autoencoder Framework for Interpreting and Steering Instruction Following of Language Models

  • 通过稀疏自编码器识别指令相关的特征激活
  • 发现特定隐变量可直接影响模型输出行为
  • 方法适配不同规模模型,适合模型可解释性研究

大型语言模型(LLM)的指令遵循能力对实际应用至关重要,但其内在机制仍不明确。本文提出一种基于稀疏自编码器(SAE)的新框架,用于解析模型如何遵循指令。通过分析SAE潜在激活,我们识别出负责指令遵循行为的特定隐变量。这些隐变量不仅在语义上与相关指令接近,还表现出对模型行为的因果影响。研究揭示了高效控制的关键因素:精确的特征定位、最后一层的作用以及指令位置优化。此外,本方法在不同大小的SAE和LLM间具有良好可扩展性。

原文摘要 · Abstract (English)

The ability of large language models (LLMs) to follow instructions is crucial for their practical applications, yet the underlying mechanisms remain poorly understood. This paper presents a novel framework that leverages sparse autoencoders (SAE) to interpret how instruction following works in these models. We demonstrate how the features we identify can effectively steer model outputs to align with given instructions. Through analysis of SAE latent activations, we identify specific latents responsible for instruction following behavior. Our findings reveal that instruction following capabilities are encoded by a distinct set of instruction-relevant SAE latents. These latents both show semantic proximity to relevant instructions and demonstrate causal effects on model behavior. Our research highlights several crucial factors for achieving effective steering performance: precise feature identification, the role of final layer, and optimal instruction positioning. Additionally, we demonstrate that our methodology scales effectively across SAEs and LLMs of varying sizes.

模型可解释性指令遵循稀疏自编码器

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。