arXiv:2503.01611cs.CL2025-03被引 3

小模型和非英语场景下,指令学习效果差,需新方法弥补差距

In-context Learning vs. Instruction Tuning: The Case of Small and Multilingual Language Models

  • 对比上下文学习与指令微调,测试多语言小模型表现
  • 非英语和小模型场景中,上下文学习性能显著下降
  • 直接偏好优化可部分提升效果,但仍不足

指令遵循是大语言模型执行下游任务的关键能力。传统指令微调依赖在精心构建的指令数据集上进行监督微调,有时还结合人类偏好对齐。近期研究显示,上下文学习(ICL)可作为替代方案引导基础模型实现指令遵循,尤其适合避免监督微调所需的大量资源。本文评估了在非英语语言和不同模型规模下,上下文学习在指令遵循任务中的可行性。结果表明,在这些场景中,上下文学习的性能明显下降。进一步实验显示,对基础模型应用直接偏好优化(DPO)可部分改善基线结果,但现有ICL方法仍无法缩小与大型英语模型之间的差距,亟需新范式。

原文摘要 · Abstract (English)

Instruction following is a critical ability for Large Language Models to perform downstream tasks. The standard approach to instruction tuning has relied on a specific phase of supervised fine-tuning over curated instruction datasets, optionally complemented with an alignment step over human preferences. Recent work has shown the potential of in-context learning (ICL) alternatives to guide base models towards instruction following. This type of approach is particularly relevant to circumvent the notable efforts and resources needed for supervised instruction tuning. In this work, we evaluate the viability of ICL for instruction following in scenarios where it is particularly relevant, i.e., languages other than English and across model sizes. Our results show that these scenarios result in downgraded ICL instruction following performance. We further show that applying Direct Preference Optimisation over base models can partially improve baseline results, although alternatives to current ICL instruction following will be needed to bridge the gap with larger English-centric language models.

指令学习小模型多语言DPO

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。