arXiv:2504.05571cs.CLcs.AI2025-04被引 10

用少量数据通过指令微调,高效注入新知识并避免遗忘。

Knowledge-Instruct: Effective Continual Pre-training from Limited Data using Instructions

  • 用合成指令数据增强小规模语料,实现知识高效注入。
  • 在新知识记忆上表现优异,同时保持原有推理能力。
  • 适合需要持续学习、数据稀缺的领域应用。

大语言模型虽在预训练中积累了大量知识,但常缺乏特定领域、新兴或小众信息。持续预训练(CPT)试图弥补这一差距,但在低数据环境下易发生灾难性遗忘且效率低下。我们提出 Knowledge-Instruct,一种通过纯指令微调从有限语料中高效注入知识的新方法。通过生成信息密集的合成指令数据,该方法能有效融入新知识,同时保留通用推理与指令遵循能力。Knowledge-Instruct 在事实记忆上表现更优,显著减少灾难性遗忘,并通过小规模语言模型生成的合成数据保持可扩展性。此外,它提升了上下文理解能力,包括复杂的多跳推理,便于与检索系统集成。我们在多个基准测试中验证其有效性,包括我们发布的 Companies——一个用于衡量知识注入能力的新数据集。

原文摘要 · Abstract (English)

While Large Language Models (LLMs) acquire vast knowledge during pre-training, they often lack domain-specific, new, or niche information. Continual pre-training (CPT) attempts to address this gap but suffers from catastrophic forgetting and inefficiencies in low-data regimes. We introduce Knowledge-Instruct, a novel approach to efficiently inject knowledge from limited corpora through pure instruction-tuning. By generating information-dense synthetic instruction data, it effectively integrates new knowledge while preserving general reasoning and instruction-following abilities. Knowledge-Instruct demonstrates superior factual memorization, minimizes catastrophic forgetting, and remains scalable by leveraging synthetic data from relatively small language models. Additionally, it enhances contextual understanding, including complex multi-hop reasoning, facilitating integration with retrieval systems. We validate its effectiveness across diverse benchmarks, including Companies, a new dataset that we release to measure knowledge injection capabilities.

持续学习指令微调知识注入小样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。