arXiv:2608.13482cs.LGcs.AI2026-08被引 1

从预训练起就植入对齐人格,让AI更早扎根人类价值观。

Synthetic Persona Pretraining: Alignment from Token Zero

论文配图:Synthetic Persona Pretraining: Alignment from Token Zero
图 1 · 摘自论文原文
  • 在预训练阶段注入价值对齐的第一人称反思,从零开始构建助手人格。
  • 30亿参数模型在5000亿词上训练后,道德困境误判率降低,抗越狱能力增强。
  • 越早介入人格塑造越有效,适合追求深层对齐的AI系统开发者。

随着基于语言模型的AI越来越多地应用于自主场景,使其目标与价值观与人类保持一致变得至关重要。目前,对齐和助手身份通常在预训练完成后才引入,这可能导致价值观成为表面附加层,而非深层根植。为此,我们提出合成人格预训练(SPP),从第一个词元起就在预训练中植入期望的助手人格。首先,我们使用规范性价值纲领生成的对齐第一人称反思标注预训练文本;其次,在标准预训练数据及其反思上通过交叉熵损失进行预训练,将期望人格嵌入多种人格之中;最后,在用户-助手对话数据上进行微调,将该人格与助手身份绑定,这一过程称为人格绑定。在5000亿词、30亿参数的模型上实验表明,SPP提升了价值纲领遵循度和越狱鲁棒性,降低了分布外道德困境中的误判率,同时保持原有能力。早期干预至关重要:相比从零开始的对齐,仅在预训练末期引入SPP会减弱价值遵循,无法改变价值优先级,并导致困境中选择更不一致。该优势依赖于人格绑定,且随预训练预算增加而提升。结果表明,早期塑造价值观对对齐至关重要,预训练阶段的人格干预是一种有效方法。

原文摘要 · Abstract (English)

As language-model-based AI is increasingly deployed in autonomous settings, aligning its goals and values with those of humans becomes critical. Today, alignment, and the assistant identity itself, are typically introduced only after pretraining, once behavioral priors are already established. This can make values a thin overlay, rather than deeply rooted, and facilitate subsequent misalignment. Pursuing a different paradigm, we introduce Synthetic Persona Pretraining (SPP), which installs the desired assistant persona from token zero in pretraining. First, we annotate pretraining documents with value-aligned first-person reflections derived from a normative value constitution. Second, we pretrain via the standard cross-entropy loss on standard pretraining documents as well as their reflections, which installs the desired persona among a multitude of other personas. Finally, we post-train on user-assistant dialogue data, which binds this desired persona to the assistant identity, a process we call persona binding. By pretraining models up to 3B parameters on 500B tokens, we show that SPP improves constitution following and jailbreak robustness, and reduces the misalignment rate in out-of-distribution moral dilemmas, while preserving capabilities. Early intervention matters: compared with alignment from token zero, introducing SPP only at the end of pretraining yields weaker constitution adherence, does not shift value priorities, and leads to less aligned choices in dilemmas. This advantage depends on persona binding and, importantly, increases with pretraining budget. Overall, our results show that shaping values early is critical for alignment and establish pretraining-time persona interventions as an effective approach to do so.

对齐预训练人格建模价值观

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。