研究大模型任务微调如何影响对齐,发现不同方法对安全等多维度有不同影响。
How LLM Task-Adaptation Reshapes Alignment: A Multi-dimensional Study of Behavioral and Representational Drift

- 对比SFT、KL-SFT和RLVR三种微调方法在15个对齐维度的表现
- SFT导致显著对齐偏移,而RLVR保持对齐更稳定,效果优于其他方法
- 揭示微调不仅是能力提升,更是对齐干预,建议纳入多维评估流程
后训练是将大语言模型适配到下游任务的关键机制。尽管已有研究指出任务适应可能改变模型原有的对齐状态,尤其是安全性行为,但其在对齐各维度的广泛影响仍不清晰。本文系统评估了代表性任务适应方法——监督微调(SFT)、KL正则化SFT和可验证奖励强化学习(RLVR)——在涵盖安全、事实性、立场稳定性、社会危害、可控性和指令遵循等六个核心领域的15个对齐维度上的表现。结果表明,后训练并非均匀重塑对齐:RLVR在提升任务性能的同时,仅引发小幅度但非零的指标特定偏移;而SFT则导致跨领域的显著对齐漂移。KL正则化可缓解此效应:更强的参考模型锚定能减少对齐漂移,但即便如此,其对齐保持能力仍不及RLVR。表示层分析进一步支持该模式,对齐相关表示的偏移与行为漂移同步。综合结果表明,任务适应不仅是能力增强步骤,本身即为一种对齐干预,呼吁将多维度对齐评估作为后训练流程的标准组成部分。
原文摘要 · Abstract (English)
Post-training is a key mechanism for adapting large language models to downstream tasks. While prior work suggests that task adaptation can alter a model's pre-existing alignment, especially its safety behavior, its broader effects across alignment domains remain poorly understood. We address this gap through a systematic evaluation of representative task-adaptation methods, including supervised fine-tuning (SFT), KL-regularized SFT, and reinforcement learning with verifiable rewards (RLVR) across 15 alignment aspects spanning six key domains: safety, factuality, stance stability, social harm, controllability, and instructability. Our results reveal that post-training does not reshape alignment uniformly. RLVR improves task performance while inducing comparatively small, but non-zero, metric-specific shifts, while SFT leads to substantially larger alignment drift across domains. KL regularization mitigates this effect: stronger reference-model anchoring reduces alignment drift from the baseline, although KL-SFT still falls short of RLVR in preserving alignment. Representation-level analysis further supports this pattern, with shifts in alignment-relevant representations tracking behavioral drift. Together, these results show that task adaptation is not merely a capability-improving step, but an alignment intervention in its own right, motivating multi-dimensional alignment evaluation as a standard component of post-training pipelines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。