arXiv:2602.15847cs.CLcs.AI2026-02ACL被引 3

发现大模型人格特征难以独立控制,存在隐性耦合。

Do Personality Traits Interfere? Geometric Limitations of Steering in Large Language Models

  • 分析五大人格特质的几何关系,发现操控一个特质会连带影响其他。
  • 即使强制正交化,仍无法完全消除跨特质影响,且削弱控制强度。
  • 适合研究可控生成、人格建模或提示工程的开发者参考。

大语言模型中的人格操控通常依赖于注入特定人格的向量,隐含假设各人格特质可独立控制。本文通过分析两种模型(LLaMA-3-8B 和 Mistral-8B)提取的五大人格操控向量,考察了其几何关系。实验采用从无约束到软/硬正交化的多种几何调控策略。结果表明,人格操控方向存在显著几何依赖:即便移除线性重叠,操控一个特质仍会引发其他特质的变化。尽管硬正交化能强制几何独立,却无法消除跨特质行为效应,且会降低操控强度。这说明大模型中的人格特质位于轻微耦合的子空间,限制了完全独立的控制能力。

原文摘要 · Abstract (English)

Personality steering in large language models (LLMs) commonly relies on injecting trait-specific steering vectors, implicitly assuming that personality traits can be controlled independently. In this work, we examine whether this assumption holds by analysing the geometric relationships between Big Five personality steering directions. We study steering vectors extracted from two model families (LLaMA-3-8B and Mistral-8B) and apply a range of geometric conditioning schemes, from unconstrained directions to soft and hard orthonormalisation. Our results show that personality steering directions exhibit substantial geometric dependence: steering one trait consistently induces changes in others, even when linear overlap is explicitly removed. While hard orthonormalisation enforces geometric independence, it does not eliminate cross-trait behavioural effects and can reduce steering strength. These findings suggest that personality traits in LLMs occupy a slightly coupled subspace, limiting fully independent trait control.

人格建模可控生成大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。