arXiv:2605.10633cs.CLcs.AI2026-05

发现语言模型的内在人格几何结构可抑制有害行为,零样本迁移有效。

Intrinsic Guardrails: How Semantic Geometry of Personality Interacts with Emergent Misalignment in LLMs

论文配图:Intrinsic Guardrails: How Semantic Geometry of Personality Interacts with Emergent Misalignment in LLMs
图 1 · 摘自论文原文
  • 用心理量表建模模型人格,发现其语义结构稳定。
  • 移除'邪恶'等向量后,错误率超40%;增强则低于3%。
  • 预训练人格向量可零样本调控微调后的有害行为。

在良性窄数据上微调大型语言模型(LLMs)有时会引发广泛有害行为,这种现象称为涌现错位(EM)。尽管已有研究将其与激活空间中的特定方向相关联,但其与模型更广泛人格表征的关系仍不清楚。本文通过大五人格、黑暗三联征及模型特有行为(如邪恶、谄媚)等心理测量学量表,映射了LLM的潜在人格空间,发现其语义几何结构在对齐模型及其受损微调版本中高度稳定。通过因果干预,我们发现隔离社会价值方向(如'邪恶'人格向量)和新提出的语义价值向量(SVV)可作为内在安全阀:删除这些向量会使错位率超过40%,而增强它们则将失败模式抑制至3%以下。基于人格空间的结构稳定性,我们进一步证明,从指令微调模型中提取的向量可零样本迁移到受损微调模型中,成功调节EM。总体表明,有害微调并未覆盖模型内部的人格表征,使得这些保守表征可作为跨分布的稳健防护机制。

原文摘要 · Abstract (English)

Fine-tuning Large Language Models (LLMs) on benign narrow data can sometimes induce broad harmful behaviors, a vulnerability termed emergent misalignment (EM). While prior work links these failures to specific directions in the activation space, their relationship to the model's broader persona remains unexplored. We map the latent personality space of LLMs through established psychometric profiles like the Big Five, Dark Triad, and LLM-specific behaviors (e.g. evil, sycophancy), and show that the semantic geometry is highly stable across aligned models and their corrupted fine-tunes. Through causal interventions, we find that directions isolating social valence, such as the 'Evil' persona vector, and a Semantic Valence Vector (SVV) that we introduce, function as intrinsic guardrails: ablating them drives the misalignment rates above $40$%, while amplifying them suppresses the failure mode to less than $3$%. Leveraging the structural stability of the personality space, we show that vectors extracted $\textit{a priori}$ from an instruct-tuned model transfer zero-shot to successfully regulate EM in corrupted fine-tunes. Overall, our findings suggest that harmful fine-tuning does not overwrite a model's internal representation of personality, allowing conserved representations to serve as robust, cross-distribution guardrails.

大模型安全人格建模鲁棒性零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。