arXiv:2512.17639cs.CL2025-12被引 11

用线性方向探测与操控大模型人格,效果因任务而异。

Linear Personality Probing and Steering in LLMs: A Big Five Study

  • 通过线性回归在激活空间中学习五大人格特质的方向。
  • 在强制选择任务中能有效操控模型行为,开放生成效果差。
  • 适合研究模型人格机制或需可控输出的场景。

大型语言模型(LLMs)展现出稳定且独特的人格特征,显著影响用户信任与参与度。现有方法或成本高(后训练),或脆弱(提示工程)。近期线性方向探测与操控成为低成本高效替代方案。本文研究是否可利用与五大人格特质对齐的线性方向实现对模型行为的探测与操控。基于 Llama 3.3 70B 模型,生成 406 个虚构角色及其五大人格评分,并使用 Alpaca 问卷的问题进行提问,采样沿人格维度变化的隐藏激活。通过线性回归学习每层的个性方向,并测试其探测与操控效果。结果表明,该方法能有效探测人格特征;但操控效果高度依赖上下文,在强制选择任务中表现可靠,而在开放生成或含额外上下文的提示中影响力有限。

原文摘要 · Abstract (English)

Large language models (LLMs) exhibit distinct and consistent personalities that greatly impact trust and engagement. While this means that personality frameworks would be highly valuable tools to characterize and control LLMs' behavior, current approaches remain either costly (post-training) or brittle (prompt engineering). Probing and steering via linear directions has recently emerged as a cheap and efficient alternative. In this paper, we investigate whether linear directions aligned with the Big Five personality traits can be used for probing and steering model behavior. Using Llama 3.3 70B, we generate descriptions of 406 fictional characters and their Big Five trait scores. We then prompt the model with these descriptions and questions from the Alpaca questionnaire, allowing us to sample hidden activations that vary along personality traits in known, quantifiable ways. Using linear regression, we learn a set of per-layer directions in activation space, and test their effectiveness for probing and steering model behavior. Our results suggest that linear directions aligned with trait-scores are effective probes for personality detection, while their steering capabilities strongly depend on context, producing reliable effects in forced-choice tasks but limited influence in open-ended generation or when additional context is present in the prompt.

大模型人格线性操控五大人格

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。