用数学方法分离大模型人格与推理能力,实现精准可控的个性化。
The Geometry of Persona: Disentangling Personality from Reasoning in Large Language Models
- 基于线性表征假设,将人格建模为正交子空间,不修改主干权重。
- 人格识别误差仅0.011,支持零样本注入且不影响原始推理能力。
- 通过向量运算可确定性控制行为,适合需要安全个性化的场景。
当前个性化大语言模型受稳定-可塑性困境制约。主流对齐方法如监督微调依赖随机权重更新,常导致‘对齐代价’——削弱通用推理能力。本文提出基于线性表征假设的Soul Engine框架,认为人格特质存在于正交线性子空间中。构建SoulBench数据集,采用冻结的Qwen-2.5基础模型与双头架构,无损提取解耦人格向量。实验显示三大突破:第一,高精度人格刻画,均方误差达0.011;第二,几何正交性验证,t-SNE可视化表明人格流形清晰连续,支持零样本人格注入且保留原模型智能;第三,确定性控制,通过向量运算实现鲁棒行为调控,经多组消融实证。本工作挑战了个性化必须微调的固有认知,从概率提示转向确定性潜空间干预,为安全可控的人工智能个性化提供数学基础。
原文摘要 · Abstract (English)
Background: The deployment of personalized Large Language Models (LLMs) is currently constrained by the stability-plasticity dilemma. Prevailing alignment methods, such as Supervised Fine-Tuning (SFT), rely on stochastic weight updates that often incur an "alignment tax" -- degrading general reasoning capabilities. Methods: We propose the Soul Engine, a framework based on the Linear Representation Hypothesis, which posits that personality traits exist as orthogonal linear subspaces. We introduce SoulBench, a dataset constructed via dynamic contextual sampling. Using a dual-head architecture on a frozen Qwen-2.5 base, we extract disentangled personality vectors without modifying the backbone weights. Results: Our experiments demonstrate three breakthroughs. First, High-Precision Profiling: The model achieves a Mean Squared Error (MSE) of 0.011 against psychological ground truth. Second, Geometric Orthogonality: T-SNE visualization confirms that personality manifolds are distinct and continuous, allowing for "Zero-Shot Personality Injection" that maintains original model intelligence. Third, Deterministic Steering: We achieve robust control over behavior via vector arithmetic, validated through extensive ablation studies. Conclusion: This work challenges the necessity of fine-tuning for personalization. By transitioning from probabilistic prompting to deterministic latent intervention, we provide a mathematically rigorous foundation for safe, controllable AI personalization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。