构建150万条高质量指令数据,提升模型对复杂任务的理解能力
Scaling Towards the Information Boundary of Instruction Sets: The Infinity Instruct Subject Technical Report
- 通过分层标签与演化生成,系统性扩展指令数据的覆盖与深度
- 在多个基础模型上验证,显著提升指令遵循能力,优于现有合成数据集
- 适合研究数据构建、模型训练优化的从业者参考
指令微调已成为释放大规模预训练模型潜力、提升其在复杂任务表现的关键。因此,构建高质量指令数据集对增强模型性能与泛化能力至关重要。尽管当前指令数据集已达数千万样本,但模型在复杂指令理解和罕见领域任务上仍存在困难,主要源于指令集在任务类型与知识领域的覆盖范围(覆盖率)以及指令复杂度(深度)上的局限。为此,我们提出一套系统性指令数据构建框架,包含分层标签体系、信息性种子选择算法、演化式数据合成流程及基于模型缺陷诊断的定向数据生成。各组件构成迭代闭环,持续提升指令数据的覆盖与深度。基于该框架,我们构建了包含约150万条指令的高质量数据集Infinity Instruct Subject。在多个基础模型与基准任务上的实验表明其有效提升指令遵循能力。进一步分析显示,相较于同类合成数据集,该数据集在覆盖范围与深度上均有显著扩展。本工作为指令数据集的高效、持续演化提供了理论与实践基础,推动从数量扩张迈向质量提升。
原文摘要 · Abstract (English)
Instruction tuning has become a foundation for unlocking the capabilities of large-scale pretrained models and improving their performance on complex tasks. Thus, the construction of high-quality instruction datasets is crucial for enhancing model performance and generalizability. Although current instruction datasets have reached tens of millions of samples, models finetuned on them may still struggle with complex instruction following and tasks in rare domains. This is primarily due to limited expansion in both ``coverage'' (coverage of task types and knowledge areas) and ``depth'' (instruction complexity) of the instruction set. To address this issue, we propose a systematic instruction data construction framework, which integrates a hierarchical tagging system, an informative seed selection algorithm, an evolutionary data synthesis process, and a model deficiency diagnosis with targeted data generation. These components form an iterative closed-loop to continuously enhance the coverage and depth of instruction data. Based on this framework, we construct Infinity Instruct Subject, a high-quality dataset containing $\sim$1.5 million instructions. Experiments on multiple foundation models and benchmark tasks demonstrate its effectiveness in improving instruction-following capabilities. Further analyses suggest that Infinity Instruct Subject shows enlarged coverage and depth compared to comparable synthesized instruction datasets. Our work lays a theoretical and practical foundation for the efficient, continuous evolution of instruction datasets, moving from data quantity expansion to qualitative improvement.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。