用荣格认知功能操控大模型人格,实现可解释的多维个性控制。
The Geometry of Personality: Activation Steering with Jungian Cognitive Functions
- 以荣格八种认知功能为框架,构建人格激活调控新方法。
- 在中间层实现对八种功能的单调可控,且多维方向不可线性分解。
- 揭示人格信息集中于中层,几何关系符合理性/非理性区分。
激活调控可实现对大语言模型的可控与可解释性,但现有研究多基于五大人格特质等静态框架。本文探索将人格视为一组认知过程,并采用荣格八种认知功能进行建模。为此,我们提出一个荣格评估协议及包含2100余条角色扮演叙事的数据集。在Llama-3.1-8B上的激活调控向量提取与评估实验表明,可通过激活调控有效实现对全部八种认知功能的单调控制。分析进一步发现:1. 人格信息集中在中间变压器层;2. 调控向量呈现符合理性与非理性功能区分的结构化几何关系;3. 有效的多维调控方向无法通过单功能方向的线性组合恢复。这些发现为理解大模型激活空间中的人格表征提供了新视角,并建立了一个可解释、高效且多维的人格控制框架。
原文摘要 · Abstract (English)
Activation steering enables control and interpretation of LLMs, yet existing work primarily models personality through static trait frameworks such as the Big Five. We investigate whether personality can instead be represented and controlled as a set of cognitive processes using the eight Jungian Cognitive Functions. To this end, we introduce a framework comprising a Jungian evaluation protocol and a dataset of over 2,100 role-playing character narrations. Activation steering vector extraction and evaluation experiments on Llama-3.1-8B demonstrate effective monotonic control over all eight cognitive functions through activation steering. Beyond controllability, our analysis reveals that: 1. personality information is concentrated in middle transformer layers; 2. steering vectors exhibit structured geometric relationships consistent with distinctions between rational and irrational functions; 3. effective multi-dimensional steering directions cannot be recovered as linear combinations of single-function directions. These findings provide new insights into the representation of personality in LLM activation space and establish a framework for studying interpretable, effective, and multi-dimensional personality control.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。