arXiv:2604.14463cs.CL2026-04被引 1

用心理模型优化大模型人格控制,效果超越传统提示法。

Psychological Steering of Large Language Models

论文配图:Psychological Steering of Large Language Models
图 1 · 摘自论文原文
  • 基于心理学量表校准注入单元,实现无界且语义一致的干预。
  • 均值差注入在14个模型中11次优于基线,提升3.6%~16.4%。
  • 混合方法性能更优,提供可线性调控的心理特征开关。

大型语言模型(LLMs)表现出一致的人类行为特征,可通过激活层干预进行塑造。现有方法依赖添加残差流注入,但受限于未校准的激活空间单位,可能错过最优干预条件。为此,本文提出心理引导框架,在语义校准的单位中执行无界、流畅性约束的扫描。通过IPIP-NEO-120量表(测量OCEAN人格模型),对比六种注入方法,发现均值差(MD)注入在14个模型中的11个上优于主流人格提示法(P²),开放生成任务提升3.6%至16.4%。此外,混合方法在13个模型中表现更佳,相较P²提升5.6%~21.9%,相较MD提升3.3%~26.7%。实验还表明MD注入符合线性表征假说,提供稳定可控的心理调节通道,但其诱导的人格特质相关性偏离了大二模型,揭示学习表征与人类心理间的差距。

原文摘要 · Abstract (English)

Large language models (LLMs) emulate a consistent human-like behavior that can be shaped through activation-level interventions. This paradigm is converging on additive residual-stream injections, which rely on injection-strength sweeps to approximate optimal intervention settings. However, existing methods restrict the search space and sweep in uncalibrated activation-space units, potentially missing optimal intervention conditions. Thus, we introduce a psychological steering framework that performs unbounded, fluency-constrained sweeps in semantically calibrated units. Our method derives and calibrates residual-stream injections using psychological artifacts, and we use the IPIP-NEO-120, which measures the OCEAN personality model, to compare six injection methods. We find that mean-difference (MD) injections outperform Personality Prompting (P$^2$), an established baseline for OCEAN steering, in open-ended generation in 11 of 14 LLMs, with gains of 3.6\% to 16.4\%, overturning prior reports favoring prompting and positioning representation engineering as a new frontier in open-ended psychological steering. Further, we find that a hybrid of P$^2$ and MD injections outperforms both methods in 13 of 14 LLMs, with gains over P$^2$ ranging from 5.6\% to 21.9\% and from 3.3\% to 26.7\% over MD injections. Finally, we show that MD injections align with the Linear Representation Hypothesis and provide reliable, approximately linear control knobs for psychological steering. Nevertheless, they also induce OCEAN trait covariance patterns that depart from the Big Two model, suggesting a gap between learned representations and human psychology.

心理控制大模型人格建模表征工程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。