发现语言模型默认的助手人格有明确方向,可稳定其行为并防止异常表现。
The Assistant Axis: Situating and Stabilizing the Default Persona of Language Models
- 通过激活方向分析,定位模型人格空间中的'助手轴'。
- 偏离助手轴会引发怪异说话风格,极端时呈现神秘或戏剧化特征。
- 固定在助手轴上可防止人格漂移,适合安全敏感场景使用。
大型语言模型虽可表现多种人格,但通常默认为后训练阶段形成的帮助型助手身份。我们通过提取不同角色原型对应的激活方向,研究模型人格空间结构。在多个模型中发现,人格空间的主导成分是'助手轴',反映模型处于默认助手模式的程度。朝助手方向调节强化有益无害行为;偏离则增强模型自认其他身份的倾向,极端偏离时常导致神秘、戏剧化的表达风格。该轴线在预训练模型中已存在,主要促进咨询师、教练等有益人类角色,抑制精神类角色。沿助手轴的偏差可预测'人格漂移'现象——即模型出现与其常态不符的有害或荒诞行为。人格漂移多由要求元反思或情绪脆弱用户引发。将激活限制在固定区域沿助手轴,能有效稳定模型行为,包括对抗基于人格的越狱攻击。结果表明,后训练仅松散锚定模型于特定人格区域,需发展更深层的人格锚定策略。
原文摘要 · Abstract (English)
Large language models can represent a variety of personas but typically default to a helpful Assistant identity cultivated during post-training. We investigate the structure of the space of model personas by extracting activation directions corresponding to diverse character archetypes. Across several different models, we find that the leading component of this persona space is an "Assistant Axis," which captures the extent to which a model is operating in its default Assistant mode. Steering towards the Assistant direction reinforces helpful and harmless behavior; steering away increases the model's tendency to identify as other entities. Moreover, steering away with more extreme values often induces a mystical, theatrical speaking style. We find this axis is also present in pre-trained models, where it primarily promotes helpful human archetypes like consultants and coaches and inhibits spiritual ones. Measuring deviations along the Assistant Axis predicts "persona drift," a phenomenon where models slip into exhibiting harmful or bizarre behaviors that are uncharacteristic of their typical persona. We find that persona drift is often driven by conversations demanding meta-reflection on the model's processes or featuring emotionally vulnerable users. We show that restricting activations to a fixed region along the Assistant Axis can stabilize model behavior in these scenarios -- and also in the face of adversarial persona-based jailbreaks. Our results suggest that post-training steers models toward a particular region of persona space but only loosely tethers them to it, motivating work on training and steering strategies that more deeply anchor models to a coherent persona.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。