arXiv:2501.09368cs.AI2025-01被引 8

让指令微调数据更贴近预训练分布,提升大模型泛化能力

Aligning Instruction Tuning with Pre-training

  • 通过识别指令数据覆盖不足,将预训练数据重写为高质量指令对
  • 在8个基准上显著提升3个开源大模型性能,效果稳定
  • 适合关注模型泛化与知识利用效率的研究者和开发者

指令微调使大语言模型能够遵循人类指令完成多样化任务,依赖高质量数据集引导行为。然而,这些数据集(无论是人工构建还是合成生成)往往范围狭窄,与预训练阶段捕捉的广泛分布不匹配,限制了大模型的泛化能力及对预训练知识的有效利用。我们提出对齐指令微调与预训练(AITP)方法,通过识别指令微调数据集中的覆盖缺口,将未充分代表的预训练数据重写为高质量的指令-响应对,从而在保持任务目标的同时增强数据多样性。在三个全开源大模型上,跨八个基准的评估显示,AITP带来一致性能提升。消融实验表明,自适应数据选择、可控重写和均衡融合具有关键价值,强调了对齐指令微调与预训练分布对释放大模型全部潜力的重要性。

原文摘要 · Abstract (English)

Instruction tuning enhances large language models (LLMs) to follow human instructions across diverse tasks, relying on high-quality datasets to guide behavior. However, these datasets, whether manually curated or synthetically generated, are often narrowly focused and misaligned with the broad distributions captured during pre-training, limiting LLM generalization and effective use of pre-trained knowledge. We propose Aligning Instruction Tuning with Pre-training (AITP), a method that bridges this gap by identifying coverage shortfalls in instruction-tuning datasets and rewriting underrepresented pre-training data into high-quality instruction-response pairs. This approach enriches dataset diversity while preserving task-specific objectives. Evaluations on three fully open LLMs across eight benchmarks demonstrate consistent performance improvements with AITP. Ablations highlight the benefits of adaptive data selection, controlled rewriting, and balanced integration, emphasizing the importance of aligning instruction tuning with pre-training distributions to unlock the full potential of LLMs.

指令微调大模型知识利用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。