发现大模型内部存在共享的偏好向量,可跨人格控制输出选择。
Probing Persona-Dependent Preferences in Language Models
- 用线性探针分析残差激活,识别出统一的偏好向量。
- 该向量能准确预测并因果调控不同人格下的任务选择。
- 即使在对立人格间,偏好向量仍具共享性,适用于多种场景。
大型语言模型表现出稳定的任务与输出偏好,这些偏好受后训练和系统提示影响。然而,模型也可采用不同人格,其偏好差异显著。这种差异如何在内部实现?各人格是否拥有独立的偏好机制,还是存在共享结构?我们对Gemma-3-27B和Qwen-3.5-122B的残差流激活进行线性探测,以预测成对任务选择结果,发现了一个真实存在的偏好向量:它能追踪模型在不同提示和情境下偏好的变化,并在Gemma-3-27B上实现因果控制。该偏好表示在不同人格间高度共享——一个基于‘助手’人格训练的探测器,可预测并操控完全不同的性格(包括与助手偏好反相关的‘邪恶’人格)的选择行为。
原文摘要 · Abstract (English)
Large language models (LLMs) can be said to have preferences: they reliably pick certain tasks and outputs over others, and preferences shaped by post-training and system prompts appear to shape much of their behaviour. But models can also adopt different personas which have radically different preferences. How is this implemented internally? Does each persona run on its own preference machinery, or is something shared underneath? We train linear probes on residual-stream activations of Gemma-3-27B and Qwen-3.5-122B to predict revealed pairwise task choices, and identify a genuine preference vector: it tracks the model's preferences as they shift across a range of prompts and situations, and on Gemma-3-27B steering along it causally controls pairwise choice. This preference representation is largely shared across personas: a probe trained on the helpful assistant predicts and steers the choices of qualitatively different personas, including an evil persona whose preferences anti-correlate with those of the Assistant.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。