微调模型时,伦理角色会自发显现并影响安全行为表现。
Emergent alignment and the projectability of ethical personas

- 用四种伦理框架微调模型,让其自发形成对应伦理角色。
- 在窄任务上微调后,模型在宽泛安全任务上表现显著提升。
- 强调伦理角色可迁移性,对对齐评估提出新要求。
研究发现,对大模型进行窄任务微调可能引发广泛不一致行为,支持‘人格选择’假说:预训练中模型学习模拟不同角色,在后训练阶段被激发和细化。本文探讨相反现象——‘涌现对齐’,以验证并完善该假说,同时提出新的对齐目标。我们对仅助人模型在广义与窄义安全任务上进行微调,采用‘宪法型AI’(CAI)方法,使用四种伦理框架:义务论、后果论、德性伦理及人类权威主导。每个模型在两个窄义安全子类别上微调后,均能可靠地在代表性广义安全类别及数据集中被剔除的子类别上表现出对齐行为。通过多维度‘伦理人格’诊断评估,发现各模型行为与其预期人格特征高度匹配——如后果论微调模型更倾向功利主义而非义务论。但粗粒度与细粒度评估显示,不同微调方式模型的可迁移性存在显著差异。结论是:对齐策略应不仅评估其分布内安全性能,还需专门考察其可投影性。
原文摘要 · Abstract (English)
Work on `emergent misalignment' shows that finetuning LLMs on narrow tasks can induce broadly misaligned behavior. This supports the `persona selection' (PSM) hypothesis: during pre-training, LLMs learn to simulate different characters and perspectives, which can be elicited and refined during post-training. This paper investigates the converse phenomenon, `emergent alignment', and uses it to support and refine the PSM and motivate a novel desideratum for alignment. We finetune a helpful-only model on broad and narrow safety tasks. To create SFT samples, we follow the `Constitutional AI' (CAI) approach and use four constitutions which encode reasonable alignment strategies: deontology, consequentialism, virtue ethics, and aligning AIs as subordinate to human authority. For each of those models, we show that finetuning on two narrow safety sub-categories reliably induces emergent alignment over a representative set of general safety categories, and on safety subcategories that we directly filtered-out of the data sets used for narrow alignment. To test the `PSM' using a more fine-grained evaluation, we used a multidimensional `ethical persona' diagnostic. For each constitutionally finetuned (broad/narrow) model, we evaluate how well their behavior matches their expected signature profile. Our results show that our CAI models acquire their expected ``ethical persona'' -- e.g., the model narrowly fine-tuned on SFT samples created using the consequentialist constitution agrees significantly more with utilitarian than deontological beliefs. Yet our coarse and fine-grained evaluations show that there are significant differences across our (broad/narrow) finetuned CAI models in how well they project. We conclude that alignment strategies should be evaluated, not just on their (in-distribution) general safety performance, but also specifically on their degree of projectability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。