arXiv:2502.11419cs.CL2025-02EMNLP

构建可持续更新的指令数据仓库,提升大模型对齐效率

InsBank: Evolving Instruction Subset for Ongoing Alignment

  • 设计渐进式数据筛选框架,动态优化指令数据质量与多样性
  • 在有限预算下生成性能更优的指令子集,显著优于基线方法
  • 适合关注大模型持续对齐与数据高效训练的研究者

大语言模型通常通过指令微调来增强对齐效果。近期研究强调指令数据的质量与多样性比数量更重要,因此需要精选多样且高质量的数据子集以降低训练成本。然而,如何随着新指令数据的产生持续演化已选数据子集仍缺乏深入探索。为此,我们提出指令银行(InsBank),一个持续更新的高质量指令数据仓库。进一步提出渐进式指令银行演化(PIBE)框架,通过基于表示的多样性评分和历史信息保留机制,实现对数据的长期高效演化。该框架支持灵活融合多样性和质量评分,在不同预算下均能提取出性能优异的子集。大量实验表明,PIBE在数据演化上显著优于基线方法,具备高度有效性和适应性。

原文摘要 · Abstract (English)

Large language models (LLMs) typically undergo instruction tuning to enhance alignment. Recent studies emphasize that quality and diversity of instruction data are more crucial than quantity, highlighting the need to select diverse, high-quality subsets to reduce training costs. However, how to evolve these selected subsets alongside the development of new instruction data remains insufficiently explored. To achieve LLMs' ongoing alignment, we introduce Instruction Bank (\textbf{InsBank}), a continuously updated repository that integrates the latest valuable instruction data. We further propose Progressive Instruction Bank Evolution (\textbf{PIBE}), a novel framework designed to evolve InsBank effectively and efficiently over time. PIBE employs a gradual data selection strategy to maintain long-term efficiency, leveraging a representation-based diversity score to capture relationships between data points and retain historical information for comprehensive diversity evaluation. This also allows for flexible combination of diversity and quality scores during data selection and ranking. Extensive experiments demonstrate that PIBE significantly outperforms baselines in InsBank evolution and is able to extract budget-specific subsets, demonstrating its effectiveness and adaptability.

指令微调数据演化大模型对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。