arXiv:2506.17289cs.AIcs.LG2025-06中稿 · ICML被引 2

对比小模型在提示与微调下的泛化能力,发现不同方法对分布外数据表现差异显著。

Evaluating Generalization and Representation Stability in Small LMs via Prompting, Fine-Tuning and Out-of-Distribution Prompts

  • 比较提示和微调在不同任务、提示风格和模型规模下的表现
  • 小模型在分布外设置下微调后性能更稳定,提示法易受数据偏移影响
  • 揭示了不同方法对任务特征抽象能力的差异,适合低数据场景选型参考

我们研究了小语言模型在两种主流适配范式——少样本提示和监督微调——下的泛化能力。尽管提示因其参数高效和灵活常被青睐,但在低资源环境和分布外变化下其鲁棒性仍不明确。本文通过对比提示与微调在不同任务格式、提示风格和模型规模下的表现,重点考察其在分布内和分布外(OOD)设置下的行为。除准确率外,还分析了各方法学习到的内部表示,评估任务特异性特征的稳定性和抽象程度。研究发现,小模型在不同适配策略下对知识的内化与泛化方式存在关键差异。该工作为低数据场景的模型选择提供了实践指导,并为提示与微调之争提供了实证见解。实验代码可在以下地址获取。

原文摘要 · Abstract (English)

We investigate the generalization capabilities of small language models under two popular adaptation paradigms: few-shot prompting and supervised fine-tuning. While prompting is often favored for its parameter efficiency and flexibility, it remains unclear how robust this approach is in low-resource settings and under distributional shifts. This paper presents a comparative study of prompting and fine-tuning across task formats, prompt styles, and model scales, with a focus on their behavior in both in-distribution and out-of-distribution (OOD) settings. Beyond accuracy, we analyze the internal representations learned by each approach to assess the stability and abstraction of task-specific features. Our findings highlight critical differences in how small models internalize and generalize knowledge under different adaptation strategies. This work offers practical guidance for model selection in low-data regimes and contributes empirical insight into the ongoing debate over prompting versus fine-tuning. Code for the experiments is available at the following

小模型提示工程泛化能力微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。