通过操控隐空间特征,揭示大模型人格形成的内在机制。
Exploring the Personality Traits of LLMs through Latent Features Steering
- 不需重训练,直接调节隐层特征以改变模型人格表现
- 发现文化规范与环境压力会影响模型人格的稳定性与一致性
- 为提升模型安全性提供基于人格特质的新视角
大型语言模型(LLMs)凭借生成类人文本的能力,显著推动了对话系统和角色扮演代理的发展。尽管已有研究显示LLMs能展现出独特且稳定的个性特征,但其编码与表达特定人格特质的机制仍不清楚。为此,我们基于社会决定论理论框架,探究文化规范与环境压力等因子如何影响模型人格。受大模型可解释性研究启发,提出一种无需训练的方法,通过提取并操控模型中对应于这些因子的隐空间特征,实现行为调控。此外,我们分析了这些因子对模型安全性的潜在影响,重点关注其通过人格维度的作用路径。
原文摘要 · Abstract (English)
Large language models (LLMs) have significantly advanced dialogue systems and role-playing agents through their ability to generate human-like text. While prior studies have shown that LLMs can exhibit distinct and consistent personalities, the mechanisms through which these models encode and express specific personality traits remain poorly understood. To address this, we investigate how various factors, such as cultural norms and environmental stressors, encoded within LLMs, shape their personality traits, guided by the theoretical framework of social determinism. Inspired by related work on LLM interpretability, we propose a training-free approach to modify the model's behavior by extracting and steering latent features corresponding to factors within the model, thereby eliminating the need for retraining. Furthermore, we analyze the implications of these factors for model safety, focusing on their impact through the lens of personality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。