arXiv:2607.26173cs.LG2026-07

跨领域移植SFT经验,提升模型对齐与鲁棒性

Shared SFT Lessons Across Alignment, Model Organisms, and Toy Models

论文配图:Shared SFT Lessons Across Alignment, Model Organisms, and Toy Models
图 1 · 摘自论文原文
  • 将对齐训练中的行为泛化方法迁移到玩具模型
  • 用非目标模型数据混合训练可保能力不降
  • 能力保留≠鲁棒,后续微调会抹除对齐行为

对齐训练、模型生物和玩具模型通常被视为独立研究方向,但三者常通过监督微调(SFT)实现相似目标。本文测试三类跨领域知识迁移:一、在玩具模型中引入对行为原因的训练,相比仅用行为示例,显著提升泛化能力;二、在模型-特定中段对齐设置中,使用非目标模型生成的数据进行SFT会损害模型能力,但加入良性同模型数据可有效避免损失;三、发现后续良性SFT会消除对齐行为却保留能力,表明能力保存并不保证对后续训练的鲁棒性。研究证明跨领域共享SFT经验可推动各方向共同进步,建议更多研究者借鉴外部方法。

原文摘要 · Abstract (English)

Alignment training, model organisms, and toy models are usually treated as separate research areas. But projects in all three frequently use supervised fine-tuning (SFT) to pursue the same underlying goals. When projects share a goal, we should test whether lessons learned from one area transfer to the other areas. We study three such transfers, each taking a lesson developed in one SFT setting and testing it in another. First, we port a lesson about behavior generalization from alignment training into toy models. Training on the reason for a behavior, as in Teaching Claude Why, can make the behavior generalize better than training on examples of the behavior alone. Second, we port a lesson about capability preservation from model organisms into the Model-Spec Midtraining alignment setting. SFT on outputs written by a model other than the student (off-model outputs) can damage capabilities when trained on. Mixing in benign on-model (and on-policy) data into our training can prevent most of this damage while still embedding the target behavior. Third, we port a lesson about robustness from model organisms into the same alignment setting. We find that follow-up benign SFT can erase the alignment behavior while preserving capabilities, showing that capability preservation alone does not ensure robustness to subsequent training. Our work illustrates how porting SFT lessons between different research fields can uplift them all, suggesting more researchers should borrow techniques from outside their own areas.

SFT对齐训练能力保留跨领域迁移

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。