通过注意力引导提升大模型概念操控效率与精度
Efficient and accurate steering of Large Language Models through attention-guided feature learning
- 基于注意力机制自动筛选相关词元嵌入,精准提取概念特征
- 在512个语义概念上实现近两倍于前代方法的成功率
- 适用于700亿参数以上模型,适合工业级大模型调优
概念操控(Steering)即直接干预大语言模型内部激活以引导其响应特定语义概念,正成为理解语义表征机制和增强模型能力的重要方向。然而现有方法极为脆弱,概念是否可操控常受特征提取算法细微调整影响。本文提出一种注意力引导的操控框架,解决三大挑战:(1) 自动选择相关词元嵌入以提取概念特征;(2) 考虑不同层中概念特征的异质性;(3) 识别最相关的操控层。在包含512个语义概念的基准测试中,该框架显著优于先前最先进方法(成功操控概念数几乎翻倍),且适用于多种模型架构与规模(最高达700亿参数)。此外,我们利用该框架揭示了概念特征在模型各层中的分布规律。整体上,该框架为开发高效、可扩展的工业级大模型微调算法开辟了新路径。
原文摘要 · Abstract (English)
Steering, or direct manipulation of internal activations to guide LLM responses toward specific semantic concepts, is emerging as a promising avenue for both understanding how semantic concepts are stored within LLMs and advancing LLM capabilities. Yet, existing steering methods are remarkably brittle, with seemingly non-steerable concepts becoming completely steerable based on subtle algorithmic choices in how concept-related features are extracted. In this work, we introduce an attention-guided steering framework that overcomes three core challenges associated with steering: (1) automatic selection of relevant token embeddings for extracting concept-related features; (2) accounting for heterogeneity of concept-related features across LLM activations; and (3) identification of layers most relevant for steering. Across a steering benchmark of 512 semantic concepts, our framework substantially improved steering over previous state-of-the-art (nearly doubling the number of successfully steered concepts) across model architectures and sizes (up to 70 billion parameter models). Furthermore, we use our framework to shed light on the distribution of concept-specific features across LLM layers. Overall, our framework opens further avenues for developing efficient, highly-scalable fine-tuning algorithms for industry-scale LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。