发现大模型人格向量在训练早期就形成并持续优化。
Tracing Persona Vectors Through LLM Pretraining

- 通过追踪训练过程,发现人格向量在0.22%预训练时已成型。
- 即使模型完全训练后,这些向量仍可有效引导行为。
- 不同提取方法揭示人格不同侧面,适用于安全可控研究。
大型语言模型如何内部表征高层次行为是影响AI安全的核心可解释性问题:它决定了我们能检测、审计或干预的内容。近期研究表明,邪恶或奉承等特质对应于内部激活中的线性方向,即人格向量。尽管这些向量现已被广泛用于安全相关的行为检测与调控,但它们在训练过程中如何形成仍不清楚。为填补这一空白,我们追踪了OLMo-3-7B在预训练阶段的人格向量,发现其在仅完成0.22%预训练时便已显著形成,并在全量微调后的指令模型中仍具有效力。虽然核心表征在早期形成,但人格向量在整个预训练过程中持续进行几何与语义上的精细化。我们还比较了多种提取策略,发现所有方法均产生有效方向,且每种策略揭示了人格的定性不同方面。在Apertus-8B上的复现表明,上述发现可定性迁移至其他模型。结果表明,人格表示是早期预训练的稳定特征,为研究训练如何形成、精炼和塑造这些表示开辟了路径。
原文摘要 · Abstract (English)
How large language models internally represent high-level behaviors is a core interpretability question with direct relevance to AI safety: it determines what we can detect, audit, or intervene on. Recent work has shown that traits such as evil or sycophancy correspond to linear directions in the internal activations, the so-called persona vectors. Although these vectors are now routinely utilized to inspect and steer model behavior in safety-relevant settings, how these representations are formed during training remains unknown. To address this gap, we trace persona vectors across the pretraining of OLMo-3-7B, finding that persona vectors form remarkably early -- within 0.22% of OLMo-3 pretraining -- and remain effective for steering the fully post-trained instruct models. Although core representations are formed early on, persona vectors continue to refine geometrically and semantically throughout pretraining. We further compare alternative elicitation strategies and find that all yield effective directions, with each strategy surfacing qualitatively distinct facets of the underlying persona. Replicating our analysis on Apertus-8B reveals that our findings transfer qualitatively beyond OLMo-3. Our results establish persona representations as stable features of early pretraining and open a path to studying how training forms, refines, and shapes them.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。