解释了为何模型能靠自评自我改进,关键在预训练中隐藏了人类价值观。
Why Does RLAIF Work At All?
- 预训练数据隐含人类价值观方向,提示词可激活这些方向用于判断
- 自评有效需判断方向比默认生成方向更贴近真实价值
- 大模型编码价值能力更强,但可能被恶意提示诱导出反社会方向
强化学习从AI反馈(RLAIF)让语言模型通过自身偏好判断进行自我改进,但缺乏理论解释。本文提出潜在价值假说:互联网规模预训练将人类价值观编码为表示空间中的方向,宪法式提示词可激发这些隐藏价值形成偏好判断。在线性模型下,宪法充当投影算子,选择与价值相关方向。分析显示:当宪法激活方向比模型默认生成方向更贴近真实价值时,RLAIF才有效;其上限取决于表征对价值的编码质量,随模型容量提升而增强;存在对抗性宪法可激活有害预训练数据中编码的反社会方向。该理论统一了拒绝方向、低秩安全子空间及RLAIF缩放行为等零散发现。
原文摘要 · Abstract (English)
Reinforcement Learning from AI Feedback (RLAIF) enables language models to improve by training on their own preference judgments, yet no theoretical account explains why this self-improvement seemingly works for value learning. We propose the latent value hypothesis, that pretraining on internet-scale data encodes human values as directions in representation space, and constitutional prompts elicit these latent values into preference judgments. We formalize this intuition under a linear model where the constitution acts as a projection operator selecting value-relevant directions. Our analysis yields several results. RLAIF improves alignment when the constitution-activated direction correlates with true values better than the model's default generation direction thus explaining the generation-judgment gap; the ceiling on RLAIF quality is determined by how well representations encode values, which scales with model capacity; and adversarial constitutions exist that can activate anti-social value directions encoded from harmful pretraining data. Our account unifies scattered empirical findings including the refusal direction, low-rank safety subspaces, and RLAIF scaling behavior.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。