arXiv:2511.03738cs.CL2025-11Conference of the …被引 4

通过提取特定层激活,实现对大模型人格特质的稳定精准控制。

Activation-Space Personality Steering: Hybrid Layer Selection for Stable Trait Control in LLMs

  • 从Transformer各层提取激活值,识别与五大人格特质对应的最佳层。
  • 发现人格特质存在于低秩共享子空间,可无损扰动实现精准引导。
  • 动态选层框架支持灵活调节,适合需要可控人格输出的场景。

大型语言模型在生成中表现出隐含人格特征,但可靠控制或对齐这些特质以满足特定需求仍是一个开放挑战。本文提出一种新流程,利用五大人格特质(开放性、尽责性、外向性、宜人性、神经质)作为心理构念,从Transformer层中提取隐藏状态激活值,结合低秩子空间发现方法,识别不同模型架构下对应人格特质的最优层,实现鲁棒注入。所得人格对齐方向通过可动态选择层的灵活引导框架进行操作化,从而精确控制大模型输出中的人格表达。研究发现人格特质占据低秩共享子空间,通过精心设计的扰动即可转化为有效引导机制,且不损害流畅性、多样性及通用能力,有助于弥合心理学理论与模型对齐之间的鸿沟。

原文摘要 · Abstract (English)

Large Language Models exhibit implicit personalities in their generation, but reliably controlling or aligning these traits to meet specific needs remains an open challenge. The need for effective mechanisms for behavioural manipulation of the model during generation is a critical gap in the literature that needs to be fulfilled. Personality-aware LLMs hold a promising direction towards this objective. However, the relationship between these psychological constructs and their representations within LLMs remains underexplored and requires further investigation. Moreover, it is intriguing to understand and study the use of these representations to steer the models' behaviour. We propose a novel pipeline that extracts hidden state activations from transformer layers using the Big Five Personality Traits (Openness, Conscientiousness, Extraversion, Agreeableness and Neuroticism), which is a comprehensive and empirically validated framework to model human personality applies low-rank subspace discovery methods, and identifies trait-specific optimal layers across different model architectures for robust injection. The resulting personality-aligned directions are then operationalised through a flexible steering framework with dynamic layer selection, enabling precise control of trait expression in LLM outputs. Our findings reveal that personality traits occupy a low-rank shared subspace, and that these latent structures can be transformed into actionable mechanisms for effective steering through careful perturbations without impacting the fluency, variance and general capabilities, helping to bridge the gap between psychological theory and practical model alignment.

人格控制大模型对齐层选择心理建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。