arXiv:2409.15825cs.CLcs.AI2024-09被引 4

60个样本即可让大模型学会问答,无需大量数据。

60 Data Points are Sufficient to Fine-Tune LLMs for Question-Answering

  • 用60个样本就能激活预训练大模型的知识
  • 不同模型需搭配不同记忆水平的数据集
  • 适合小样本场景下的模型微调研究

大型语言模型(LLMs)通过海量数据预训练获得广泛世界知识,可进一步微调用于问答任务。然而,针对该任务的有效微调策略仍不明确。为此,我们根据预训练模型对知识的存储程度对监督微调(SFT)数据进行分类,并开展了一系列实证分析。实验涵盖三个模型家族中的四款大模型,重点考察三个关键因素:SFT所需的数据量、不同SFT数据集对模型性能的影响,以及数据需求在不同模型间的差异。结果表明,仅需60个数据点即可激活预训练阶段编码的知识,使模型具备问答能力。此外,不同记忆水平的数据对模型表现有显著影响,最优数据集随具体模型而异。未来研究将深入探讨其内在机制。

原文摘要 · Abstract (English)

Large language models (LLMs) encode extensive world knowledge through pre-training on massive datasets, which can then be fine-tuned for the question-answering (QA) task. However, effective strategies for fine-tuning LLMs for the QA task remain largely unexplored. To address this gap, we categorize supervised fine-tuning (SFT) data based on the extent of knowledge memorized by the pretrained LLMs and conduct a series of empirical analyses. Our experiments, involving four LLMs from three different model families, focus on three key factors: the amount of data required for SFT, the impact of different SFT datasets on model performance, and how data requirements vary across LLMs. The results show that as few as 60 data points during the SFT stage can activate the knowledge encoded during pre-training, enabling LLMs to perform the QA task. Additionally, SFT with data of varying memory levels has a significant impact on LLM performance, with the optimal dataset differing based on the specific model being fine-tuned. Future research will delve deeper into the mechanisms underlying these phenomena.

大模型微调小样本学习问答系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。