arXiv:2410.09629cs.CLcs.AI2024-10EMNLP被引 11

用合成数据提升大模型知识注入与提炼能力

Synthetic Knowledge Ingestion: Towards Knowledge Refinement and Injection for Enhancing Large Language Models

  • 通过细粒度合成与交替生成构建高质量知识数据
  • 在金融、生物医学等领域显著优于基线方法
  • 适合需要精准知识更新的场景,如专业问答

大语言模型虽能掌握多领域事实知识,但在已有知识的优化或外部知识的整合方面仍面临挑战。本文提出一种名为Ski的新型合成知识注入方法,采用细粒度合成、交错生成和集成增强策略,从原始知识源构建高质量数据表示。随后,将Ski及其变体与三种知识注入技术——检索增强生成(RAG)、监督微调(SFT)和持续预训练(CPT)——结合,用于模型中的知识注入与精炼。在涵盖金融、生物医学及开放生成领域的多个问答任务上开展大量实证实验,结果表明Ski显著优于基线方法,有效提升了知识注入效果。本工作为提升大模型输出的事实准确性提供了重要路径,推动了知识表征与注入能力的改进。

原文摘要 · Abstract (English)

Large language models (LLMs) are proficient in capturing factual knowledge across various domains. However, refining their capabilities on previously seen knowledge or integrating new knowledge from external sources remains a significant challenge. In this work, we propose a novel synthetic knowledge ingestion method called Ski, which leverages fine-grained synthesis, interleaved generation, and assemble augmentation strategies to construct high-quality data representations from raw knowledge sources. We then integrate Ski and its variations with three knowledge injection techniques: Retrieval Augmented Generation (RAG), Supervised Fine-tuning (SFT), and Continual Pre-training (CPT) to inject and refine knowledge in language models. Extensive empirical experiments are conducted on various question-answering tasks spanning finance, biomedicine, and open-generation domains to demonstrate that Ski significantly outperforms baseline methods by facilitating effective knowledge injection. We believe that our work is an important step towards enhancing the factual accuracy of LLM outputs by refining knowledge representation and injection capabilities.

知识注入大模型优化合成数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。