arXiv:2604.11802cs.CL2026-04

发现大五人格神经元,可精准操控模型性格表现

Psychological Concept Neurons: Can Neural Control Bias Probing and Shift Generation in LLMs?

  • 定位大五人格在模型中选择性激活的神经元
  • 干预神经元使性格探测准确率超0.8,定向控制有效
  • 适合研究模型心理表征与可控生成的学者

利用大五人格等心理学构念,大型语言模型(LLMs)可模拟特定人格特征并预测用户人格。尽管模型行为与这些构念一致,但其内部表示位置与机制尚不明确。本文聚焦问卷操作化的大五人格概念,分析其内部表征的形成与定位,并通过干预探究其与行为输出的关系。实验发现,大五信息在早期层即可快速解码,且贯穿全模型;概念选择性神经元主要集中在中层,跨领域重叠有限。对这些神经元进行增强或抑制,能持续将探测读数偏向目标概念,部分概念成功率达0.8以上,表明模型内部人格特质可被因果调控。在标签生成层面,相同干预虽常引发预期方向偏移,但效果较弱、依赖具体概念,且常伴随跨特质溢出,显示生成行为控制仍具挑战。整体揭示了模型表征控制与行为控制之间的差距。

原文摘要 · Abstract (English)

Using psychological constructs such as the Big Five, large language models (LLMs) can imitate specific personality profiles and predict a user's personality. While LLMs can exhibit behaviors consistent with these constructs, it remains unclear where and how they are represented inside the model and how they relate to behavioral outputs. To address this gap, we focus on questionnaire-operationalized Big Five concepts, analyze the formation and localization of their internal representations, and use interventions to examine how these representations relate to behavioral outputs. In our experiment, we first use probing to examine where Big Five information emerges across model depth. We then identify neurons that respond selectively to each Big Five concept and test whether enhancing or suppressing their activations can bias latent representations and label generation in intended directions. We find that Big Five information becomes rapidly decodable in early layers and remains detectable through the final layers, while concept-selective neurons are most prevalent in mid layers and exhibit limited overlap across domains. Interventions on these neurons consistently shift probe readouts toward targeted concepts, with targeted success rates exceeding 0.8 for some concepts, indicating that the model's internal separation of Big Five personality traits can be causally steered. At the label-generation level, the same interventions often bias generated label distributions in the intended directions, but the effects are weaker, more concept-dependent, and often accompanied by cross-trait spillover, indicating that comparable control over generated labels is difficult even with interventions on a large fraction of concept-selective neurons. Overall, our findings reveal a gap between representational control and behavioral control in LLMs.

大五人格神经元干预可控生成心理表征

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。