用概念激活向量精准控制大模型输出,无需大量训练。
Controlling Large Language Models Through Concept Activation Vectors
- 通过训练概念激活向量,实现对特定概念的精确控制。
- 在毒性、情感、风格和主题控制上均达到顶尖效果。
- 轻量级设计,可对单个样本灵活调整控制强度与位置。
随着大语言模型(LLMs)在各领域的广泛应用,对其生成内容的可控性要求日益提高,需使其输出符合人类价值观或满足个性化主题与风格需求。现有方法或计算开销大,或控制粒度粗。本文提出生成用概念激活向量(GCAV),一种轻量级控制框架,可在不进行资源密集型微调的前提下实现精准控制。具体而言,GCAV首先为特定概念(如毒性)训练概念激活向量;推理时,通过在模型激活层中调整该向量(如移除毒性向量)实现控制。多角度实验(包括毒性降低、情感调节、语言风格与主题控制)表明,该框架在细粒度控制上表现优异,支持对不同样本的控制层与控制幅度灵活调整。
原文摘要 · Abstract (English)
As large language models (LLMs) are widely deployed across various domains, the ability to control their generated outputs has become more critical. This control involves aligning LLMs outputs with human values and ethical principles or customizing LLMs on specific topics or styles for individual users. Existing controlled generation methods either require significant computational resources and extensive trial-and-error or provide coarse-grained control. In this paper, we propose Generation with Concept Activation Vector (GCAV), a lightweight model control framework that ensures accurate control without requiring resource-extensive fine-tuning. Specifically, GCAV first trains a concept activation vector for specified concepts to be controlled, such as toxicity. During inference, GCAV steers the concept vector in LLMs, for example, by removing the toxicity concept vector from the activation layers. Control experiments from different perspectives, including toxicity reduction, sentiment control, linguistic style, and topic control, demonstrate that our framework achieves state-of-the-art performance with granular control, allowing for fine-grained adjustments of both the steering layers and the steering magnitudes for individual samples.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。