arXiv:2606.20814cs.AIcs.LG2026-06

研究小规模微调如何导致模型广泛且不均衡的对齐偏差。

What Shapes Emergent Misalignment? Insights from Training Dynamics, Model Priors, and Data

论文配图:What Shapes Emergent Misalignment? Insights from Training Dynamics, Model Priors, and Data
图 1 · 摘自论文原文
  • 通过分析训练动态、模型先验和数据,揭示偏差成因。
  • 预训练模型激活能预测微调后的对齐得分,存在正相关信号。
  • 训练与评估激活变化在子空间中重叠度高,提示偏差可预测。

涌现对齐偏差(EM)是指模型在小规模微调后泛化时出现广泛但不均匀的对齐问题。本文直接从微调的三个核心成分——训练动态、模型先验和数据——出发研究其成因。首先,我们发现域内训练损失与域外对齐分数之间存在关联,但不同学习率调度未能显著提升广义对齐表现。其次,尽管微调后模型的偏差得分均值与标准差通常显著不同于预训练模型,但预训练及指令模型的仅提示激活能有效预测微调后的精细对齐分数,显示潜在正相关。最后,我们比较了微调前后训练与评估提示的激活变化,发现二者在子空间中的重叠度中等至较高,且基于最后一个提示词激活的测量显示,激活偏移相似性与子空间重叠高度相关。该重叠被控制在随机向量与评估激活对比基线之上。

原文摘要 · Abstract (English)

Emergent misalignment (EM) is a phenomenon in which models generalize with narrow fine-tuning, leading to broad (yet uneven) misalignment across evaluation questions. We study EM and its variability directly through the components of fine-tuning: training dynamics, model priors, and data. (1) We first explored how in-domain training loss relates to out-of-domain alignment scores across datasets and model families. Then, we tried to induce potential alternative local minima through different learning schedules for one narrow fine-tuning, but did not find strong runs with better broad alignment scores conditioned on similar or lower training loss. (2) We found that although the mean and standard deviations of the misaligned model scores are usually statistically different from those of the pre-trained model, there are some potential signals on overall positive correlation. The evaluation prompt-only activations from both the pre-trained and the original instruct models (prior to narrow fine-tuning) could predict fine-grained alignment scores after narrow fine-tuning. (3) Finally, we compared activation deltas before and after narrow fine-tuning and found moderate-to-high subspace overlap and similarity between the resulting activation shifts for training and evaluation prompts. Subspace overlaps between training and evaluation prompt activations correlate with their shifts' similarities when measuring with the last prompt-token activations. The train-evaluation data prompt overlap is controlled against overlap computed from random vectors and evaluation prompts activations.

对齐偏差微调研究模型先验激活分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。