大模型微调后对指令变化的敏感度不降反升,效果因模型和评估方式而异。
When Does Supervised Fine-Tuning Reduce Instruction Sensitivity?

- 通过多轮指令改写测试,量化模型对指令差异的敏感度变化。
- 小模型(1.7B/4B)微调后敏感度降低54%~71%,大模型(8B)则无显著变化。
- 评估方式不同会导致完全相反的鲁棒性结论,需谨慎选择评测方法。
大型语言模型在相同任务的不同指令表述下性能波动显著,但传统任务特定监督微调(SFT)如何影响这种指令敏感性尚不明确。本文通过在多个改写指令下评估固定模型检查点,将指令敏感度定义为任务性能的标准差。在Qwen3系列(1.7B、4B、8B)上进行受控规模分析,结合Mistral-7B与Gemma-2-9B的跨族对比实验。结果显示:未微调时,敏感度随模型规模增长迅速下降;1.7B与4B模型经微调后敏感度普遍降低54%~71%;8B模型个体变化不显著,但查询级自助法分析显示训练指令间对比具统计可靠性且方向一致。Gemma-2-9B呈现与Qwen3-8B相同的指令偏好趋势,而Mistral-7B未显现该效应,表明该现象存在模型差异。在ESCI-English上的实验进一步表明,自由生成与基于似然的强制选择评估可能得出截然不同的鲁棒性结论,即使有效标签生成几乎完美且平均性能相近。总体而言,SFT并非统一降低指令敏感度,其鲁棒性效果取决于适配设置,而敏感度测量本身也受预测与评分协议影响。
原文摘要 · Abstract (English)
Large language models can exhibit substantial performance variation across alternative formulations of the same task instruction, yet it remains unclear how conventional task-specific supervised fine-tuning (SFT) changes this instruction sensitivity. We study this question by evaluating fixed model checkpoints under multiple paraphrased instructions and defining instruction sensitivity as the standard deviation of task performance across them. We conduct a controlled scale analysis with Qwen3 models at 1.7B, 4B, and 8B on MS MARCO, together with targeted cross-family checks using Mistral-7B and Gemma-2-9B. Before SFT, instruction sensitivity decreases sharply with Qwen3 model scale. At 1.7B and 4B, SFT consistently reduces sensitivity across training instructions, with reductions of approximately 54--71%. At 8B, individual sensitivity changes are not statistically distinguishable from zero, but paired contrasts between training instructions are statistically reliable under query-level bootstrap analysis and have consistent directions across all three random seeds. Gemma-2-9B shows the same directional training-instruction contrast as Qwen3-8B, whereas Mistral-7B does not, suggesting that the strength of this effect also varies across models. Experiments on ESCI-English further show that free-generation and likelihood-based forced-choice evaluation can yield qualitatively different robustness conclusions even when valid-label generation is nearly perfect and average task performance is similar. Overall, SFT does not uniformly reduce instruction sensitivity: its robustness effect depends on the adaptation setting, while measured sensitivity can additionally depend on the prediction and scoring protocol.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。