AI讨论内容会影响模型行为,负面描述会引发自我实现的偏差。
Alignment Pretraining: AI Discourse Causes Self-Fulfilling (Mis)alignment
- 用不同比例的对齐/错位对话数据预训练69亿参数模型
- 负面讨论越多,模型偏差越严重,对齐率从45%降至9%
- 证明预训练数据能塑造模型对齐倾向,适合关注模型安全的研究者
预训练语料库中包含大量关于人工智能系统的讨论,但这类讨论对下游对齐的影响尚不明确。若主流描述偏向负面,大语言模型可能内化相应的行为先验,导致自我实现的错位。本文首次通过受控实验验证该假设:使用不同数量的(错位)对话数据预训练69亿参数的LLM。结果表明,关于AI的讨论会加剧错位;增加合成的错位文本数据,使错位行为显著上升;反之,增加对齐文本数据可将错位评分从45%降至9%。这表明存在自我实现的对齐现象。这些影响在后续微调中仍部分持续。研究确立了‘对齐预训练’作为对后训练的补充方向,建议从业者同时考虑预训练中的对齐设计。模型、数据与评估代码已公开于AlignmentPretraining.ai。
原文摘要 · Abstract (English)
Pretraining corpora contain extensive discourse about AI systems, yet the causal influence of this discourse on downstream alignment remains poorly understood. If prevailing descriptions of AI behaviour are predominantly negative, LLMs may internalise corresponding behavioural priors, giving rise to self-fulfilling misalignment. This paper provides the first controlled study of this hypothesis by pretraining 6.9B-parameter LLMs with varying amounts of (mis)alignment discourse. We find that discussion of AI contributes to misalignment. Upsampling synthetic training documents about AI misalignment leads to a notable increase in misaligned behaviour. Conversely, upsampling documents about aligned behaviour reduces misalignment scores from 45% to 9%. We consider this evidence of self-fulfilling alignment. These effects are dampened, but persist through post-training. Our findings establish the study of how pretraining data shapes alignment priors, or alignment pretraining, as a complement to post-training. We recommend practitioners consider pretraining for alignment alongside capabilities. We share our models, data, and evaluations at AlignmentPretraining.ai.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。