发现不同大模型的概念表示可通过线性变换对齐,小模型的控制向量可操控大模型行为。
Cross-model Transferability among Large Language Models on the Platonic Representations of Concepts
- 用线性变换对齐不同大模型的概念表示,实现跨模型行为控制。
- 小模型提取的控制向量能有效操纵大模型生成内容。
- 该方法通用性强,适用于多种概念的跨模型迁移。
理解大语言模型(LLM)的内部机制是关键研究方向。已有研究证明,单个LLM中的概念表示可通过控制向量(SVs)捕捉,从而实现对模型行为的调控(如生成有害内容)。本文提出新视角,探索不同LLM间概念表示的内在关联,类比柏拉图洞穴寓言。我们引入线性变换方法来连接这些表示,得出三个关键发现:1)不同LLM的概念表示可通过简单线性变换有效对齐,实现高效的跨模型迁移与行为控制;2)该变换在概念间具有泛化能力,可统一调控不同概念的控制向量;3)存在弱到强的迁移能力,即从小模型提取的控制向量可有效操控大模型行为。
原文摘要 · Abstract (English)
Understanding the inner workings of Large Language Models (LLMs) is a critical research frontier. Prior research has shown that a single LLM's concept representations can be captured as steering vectors (SVs), enabling the control of LLM behavior (e.g., towards generating harmful content). Our work takes a novel approach by exploring the intricate relationships between concept representations across different LLMs, drawing an intriguing parallel to Plato's Allegory of the Cave. In particular, we introduce a linear transformation method to bridge these representations and present three key findings: 1) Concept representations across different LLMs can be effectively aligned using simple linear transformations, enabling efficient cross-model transfer and behavioral control via SVs. 2) This linear transformation generalizes across concepts, facilitating alignment and control of SVs representing different concepts across LLMs. 3) A weak-to-strong transferability exists between LLM concept representations, whereby SVs extracted from smaller LLMs can effectively control the behavior of larger LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。