无需人工参与,用迭代训练高效筛选高质量指令数据。
IterSelectTune: An Iterative Training Framework for Efficient Instruction-Tuning Data Selection
- 通过迭代训练自动筛选源数据中20%的优质指令样本。
- 仅用20%数据训练的模型在多个基准上超越全量数据微调结果。
- 适合追求低成本高效微调的AI研究者与工程师使用。
随着大语言模型(LLMs)持续发展,指令微调对提升其生成准确且符合上下文响应的能力变得至关重要。尽管已有大量指令微调数据集,但从大规模源数据集中筛选高质量指令数据通常需要大量人力。本文提出$ extbf{IterSelectTune}$,一种无需人工参与、依赖极少GPT-4的高效、低成本迭代训练策略,用于指令数据选择。通过对约20%源数据进行微调,该方法在多个基准测试和公开测试集上始终优于在全量数据上微调的模型。结果表明,该方法在降低计算资源消耗的同时有效提升了LLM性能。
原文摘要 · Abstract (English)
As large language models (LLMs) continue to advance, instruction tuning has become critical for improving their ability to generate accurate and contextually appropriate responses. Although numerous instruction-tuning datasets have been developed to enhance LLM performance, selecting high-quality instruction data from large source datasets typically demands significant human effort. In this work, we introduce $\textbf{IterSelectTune}$, an efficient, cost-effective iterative training policy for selecting high-quality instruction data with no human involvement and limited reliance on GPT-4. By fine-tuning on approximately 20\% of the source data, our method consistently outperforms models fine-tuned on the full dataset across multiple benchmarks and public test datasets. These results highlight the effectiveness of our approach in enhancing LLM performance while reducing the computational resources required for instruction tuning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。