指令多样性决定模型泛化能力,跨领域数据更有效。
$\textbf{Only-IF}$:Revealing the Decisive Effect of Instruction Diversity on Generalization
- 通过跨领域指令数据提升模型泛化能力。
- 相同数据量下,多样化指令显著提升性能。
- 适用于专家与通用模型的训练优化指导。
理解并准确遵循指令对大语言模型在多样任务中的表现至关重要。本文通过受控实验,基于图灵完备马尔可夫算法思想,证明模型泛化能力仅在训练数据具备足够语义领域多样性时才会出现。单纯在有限领域内多样化无法保证鲁棒泛化。相比之下,即使数据预算受限,跨领域数据多样化也能显著增强模型适应性。我们在真实场景中分析了专业型和通用型模型的微调,结果表明:1)在保持数据量不变的前提下,增加已有数据集的多样性可获得更好性能;2)数据扩展时,多样化指令语义比简单增加相似数据更有效。研究为指令微调数据构建提供关键洞见,尤其在扩展训练数据以优化专业型和通用型模型性能时,需注重数据的语义多样性。我们强调战略化多样性的重要性,并给出明确的数据质量提升指南。
原文摘要 · Abstract (English)
Understanding and accurately following instructions is critical for large language models (LLMs) to be effective across diverse tasks. In this work, we rigorously examine the key factors that enable models to generalize to unseen instructions, providing insights to guide the collection of data for instruction-tuning. Through controlled experiments, inspired by the Turing-complete Markov algorithm, we demonstrate that such generalization $\textbf{only emerges}$ when training data is diversified enough across semantic domains. Our findings also reveal that merely diversifying within limited domains fails to ensure robust generalization. In contrast, cross-domain data diversification, even under constrained data budgets, significantly enhances a model's adaptability. We further extend our analysis to real-world scenarios, including fine-tuning of $\textit{$\textbf{specialist}$}$ and $\textit{$\textbf{generalist}$}$ models. In both cases, we demonstrate that 1) better performance can be achieved by increasing the diversity of an established dataset while keeping the data size constant, and 2) when scaling up the data, diversifying the semantics of instructions is more effective than simply increasing the quantity of similar data. Our research provides important insights for dataset collation, particularly when optimizing model performance by expanding training data for both specialist and generalist scenarios. We show that careful consideration of data diversification is key: training specialist models with data extending beyond their core domain leads to significant performance improvements, while generalist models benefit from diverse data mixtures that enhance their overall instruction-following capabilities across a wide range of applications. Our results highlight the critical role of strategic diversification and offer clear guidelines for improving data quality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。