通过指令微调大幅提升大模型性能,支持多种任务与场景。
Scaling Instruction-Finetuned Language Models

- 在1800个任务上进行指令微调,提升模型泛化能力。
- Flan-PaLM 540B在五轮测试中达到75.2%准确率,超越原版模型9.4%。
- 释放Flan-T5模型,小体积也能媲美大模型表现。
在指令形式的数据集上对语言模型进行微调已被证明能提升模型性能和对未见任务的泛化能力。本文重点研究了三个方面的扩展:(1)任务数量的增加,(2)模型规模的扩大,(3)链式思维(chain-of-thought)数据的微调。实验表明,结合上述策略的指令微调显著提升了多种模型架构(PaLM、T5、U-PaLM)、提示设置(零样本、少样本、链式思维)及评估基准(MMLU、BBH、TyDiQA、MGSM、开放生成)上的表现。例如,基于1800个任务微调的Flan-PaLM 540B相比原始PaLM 540B平均提升9.4%。该模型在多个基准上达到顶尖水平,如五轮测试下在MMLU上达75.2%。此外,本文公开发布了Flan-T5模型,其少样本表现甚至优于更大的PaLM 62B模型。总体而言,指令微调是一种通用且高效的提升预训练语言模型性能与可用性的方法。
原文摘要 · Abstract (English)
Finetuning language models on a collection of datasets phrased as instructions has been shown to improve model performance and generalization to unseen tasks. In this paper we explore instruction finetuning with a particular focus on (1) scaling the number of tasks, (2) scaling the model size, and (3) finetuning on chain-of-thought data. We find that instruction finetuning with the above aspects dramatically improves performance on a variety of model classes (PaLM, T5, U-PaLM), prompting setups (zero-shot, few-shot, CoT), and evaluation benchmarks (MMLU, BBH, TyDiQA, MGSM, open-ended generation). For instance, Flan-PaLM 540B instruction-finetuned on 1.8K tasks outperforms PALM 540B by a large margin (+9.4% on average). Flan-PaLM 540B achieves state-of-the-art performance on several benchmarks, such as 75.2% on five-shot MMLU. We also publicly release Flan-T5 checkpoints, which achieve strong few-shot performance even compared to much larger models, such as PaLM 62B. Overall, instruction finetuning is a general method for improving the performance and usability of pretrained language models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。