用稀疏自编码器特征精准控制大模型行为,效果更好且可解释。
Interpretable Steering of Large Language Models with Feature Guided Activation Additions
- 在稀疏自编码器的潜在空间中选择特定特征构建引导向量。
- 在Gemma-2系列模型上,各项任务表现优于现有方法。
- 揭示了引导强度与模型通用能力间的权衡关系。
大语言模型的行为可控性是重要挑战。现有激活引导方法虽有潜力,但精确性和可解释性不足。本文提出特征引导激活添加(FGAA),结合对比激活添加(CAA)和稀疏自编码器靶向引导(SAE-TS)的洞察,在稀疏自编码器(SAE)的潜在空间中通过优化选择目标特征,构建更精准的引导向量,提升引导效果并保持输出连贯性。在Gemma-2-2B和Gemma-2-9B模型上,多种引导任务测试表明,FGAA显著优于CAA、SAE解码器引导及SAE-TS等方法。结果还揭示了所有方法均存在的引导尺度与模型通用能力之间的权衡关系。
原文摘要 · Abstract (English)
Effective and reliable control over large language model (LLM) behavior is a significant challenge. While activation steering methods, which add steering vectors to a model's hidden states, are a promising approach, existing techniques often lack precision and interpretability in how they influence model outputs. We introduce Feature Guided Activation Additions (FGAA), a novel activation steering method that leverages insights from Contrastive Activation Addition (CAA) and Sparse Autoencoder-Targeted Steering (SAE-TS). By operating in the latent space of a Sparse Autoencoder (SAE) and employing optimization techniques to select desired SAE features, FGAA constructs precise steering vectors that provide better steering effects while maintaining coherence of steered model outputs. In this regard, evaluations on Gemma-2-2B and Gemma-2-9B models across various steering tasks demonstrate that FGAA outperforms existing steering methods of CAA, SAE decoder steering, and SAE-TS. Our results also highlight important trade-offs between steering scale and general model capabilities that are consistent across all tested steering methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。