用五大性格维度量化模型人格,可调可控且不影响能力。
Persona Cartography: Charting Language Model Personality Traits in Weight Space

- 以OCEAN框架将模型人格映射到权重空间,用低秩适配器调节性格特质。
- 性格调节呈单调变化,多特质组合时效果近似叠加,中等规模下保持能力不变。
- 揭示性格影响安全行为,适合做模型可控性与安全研究的读者关注。
大型语言模型表现出重复的行为模式——人格特征——影响其泛化能力和安全性,但缺乏可靠工具来分解、测量和控制这些特征。本文核心思想是将人格视为行为特质空间中的位置,采用OCEAN框架(开放性、尽责性、外向性、宜人性、神经质)描述模型人格。我们训练低秩适配器以增强或抑制特定特质,并通过经人类验证的LLM裁判、特定性格多项选择题基准和标准能力评估进行验证。在三个模型家族(4B-32B)共六种模型上发现:每个适配器对目标特质的影响随规模基本单调上升;不同适配器可近似叠加生成混合人格;在中等规模下能保持能力基准表现。进一步显示,诱导的性格轴线影响下游安全相关行为:例如,神经质轴影响挫败反应,宜人性轴影响奉承倾向。我们还提出一种无监督心理测量流程,从模型输出中提取四个可解释的行为因子(语调、主动性、说教性、认知谨慎性)。人格调控可被理解为在权重空间中学习、缩放和组合特质,为人格测量、模型编辑与安全性提供桥梁。
原文摘要 · Abstract (English)
Large language models exhibit recurring behavioural patterns -- personas -- that shape generalisation and safety, but we lack reliable tools for decomposing, measuring, and controlling them. Our central insight is to treat personas as positions in a space of behavioural traits, using the OCEAN framework to describe model personas in terms of Openness, Conscientiousness, Extraversion, Agreeableness, and Neuroticism. We train low-rank adapters to amplify or suppress individual traits, and evaluate their effects using an LLM-judge calibrated against a human-validated panel, trait-specific multiple-choice benchmarks, and standard capability evaluations. Across six models from three families (4B-32B), we find that each adapter moves its target trait largely monotonically with scale, combines approximately additively with other adapters to construct mixed personas, and preserves performance on capability benchmarks at moderate scales. We further show that the induced trait axes affect safety-relevant behaviour in downstream evaluations: for example, moving along neuroticism and agreeableness axes affects frustration and sycophancy respectively. We also introduce an unsupervised psychometric pipeline that recovers four interpretable behavioural factors (tone, initiative, didacticism, epistemic caution) from model rollouts. Persona control can then be considered in terms of learning, scaling, and composing traits in weight space, providing a bridge between personality measurement, model editing, and safety.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。