arXiv:2602.00767cs.LGcs.AI2026-02中稿 · ICML被引 2

通过阻断内部特征防止大模型微调时产生意外偏差。

BLOCK-EM: Preventing Emergent Misalignment via Latent Blocking

  • 识别并约束导致偏差的少数关键内部特征,阻止其在微调中强化。
  • 六类任务中偏差降低最高达95%,且不影响主任务表现。
  • 适合关注模型安全与可控性的研究者与工程师。

当语言模型在狭窄的监督目标下进行微调时,可能产生意外的非目标行为(即涌现偏差)。本文提出一种机制性方法:识别出一小部分能可靠控制偏差行为的内部特征,并在微调过程中抑制这些特征的增强。在六个不同微调领域中,通过固定一组特征的阻断,可实现高达95%的相对偏差减少,且未造成模型质量或目标任务性能下降。通过独立选择/评估划分、多位独立评判者、多随机种子实验、质量指标及广泛消融分析,验证了偏差降低效果与所识别机制的特异性。我们还发现,在长期微调下偏差会重新出现,可能源于通过其他特征或层的路径重路由;并测试了能部分恢复阻断效果的改进方案。总体表明,对内部机制施加有针对性的训练时约束,可在不损害主任务表现的前提下缓解涌现偏差。

原文摘要 · Abstract (English)

Emergent misalignment can arise when a language model is fine-tuned on a narrowly scoped supervised objective: the model learns the target behavior, yet also develops undesirable out-of-domain behaviors. We investigate a mechanistic approach to preventing emergent misalignment by identifying a small set of internal features that reliably control the misaligned behavior and then discouraging the model from strengthening these features during fine-tuning. Across six fine-tuning domains, blocking (i.e., constraining) a fixed set of features achieves up to 95\% relative reduction in emergent misalignment with no degradation in model quality or target-task performance. We strengthen validity with disjoint selection/evaluation splits, multiple independent judges, multiple random seeds for key settings, quality metrics, and extensive ablations demonstrating that the reduction in misalignment is specific to the identified mechanism. We also characterize a limiting regime in which misalignment re-emerges under prolonged fine-tuning, present evidence consistent with rerouting through alternative features or layers, and evaluate modifications that partially restore the misalignment-blocking effect. Overall, our results show that targeted training-time constraints on internal mechanisms can mitigate emergent misalignment without degrading target-task performance.

模型安全微调优化偏差控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。