指令微调未必更好,数学和跨域任务中基模型反而更优。
Do Instruction-Tuned Models Always Perform Better Than Base Models? Evidence from Math and Domain-Shifted Benchmarks
- 在零样本思维链下,基模型比指令微调模型表现更稳。
- 指令微调模型在医学计算任务上性能下降超32.67%。
- 适合关注提示工程与模型真实推理能力的研究者。
指令微调是提升大模型性能的常用方法,但其是否真正增强推理能力仍不明确。我们评估了基模型与指令微调模型在标准数学基准、结构扰动数据集以及跨领域任务上的表现。结果揭示两个常被忽视的局限:第一,在零样本思维链(zero-shot CoT)设置下,基模型在GSM8K上始终优于指令微调版本,性能降幅高达32.67%(Llama3-70B),仅当提供少样本示例时指令模型才可追平或超越,表明其依赖特定提示模式而非内在推理;第二,指令微调模型在分布外任务中表现脆弱,于领域特定的MedCalc基准上被基模型超越,且在扰动数据集上出现显著性能下降,反映出对提示结构的敏感性而非鲁棒推理能力。
原文摘要 · Abstract (English)
Instruction finetuning is standard practice for improving LLM performance, yet it remains unclear whether it enhances reasoning or merely induces surface-level pattern matching. We investigate this by evaluating base and instruction-tuned models on standard math benchmarks, structurally perturbed variants, and domain-shifted tasks. Our analysis highlights two key (often overlooked) limitations of instruction tuning. First, the performance advantage is unstable and depends heavily on evaluation settings. In zero-shot CoT settings on GSM8K, base models consistently outperform instruction-tuned variants, with drops as high as 32.67\% (Llama3-70B). Instruction-tuned models only match or exceed this performance when provided with few-shot exemplars, suggesting a reliance on specific prompting patterns rather than intrinsic reasoning. Second, tuning gains are brittle under distribution shift. Our results show that base models surpass instruction-tuned variants on the domain-specific MedCalc benchmark. Additionally, instruction-tuned models show sharp declines on perturbed datasets, indicating sensitivity to prompt structure over robust reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。