arXiv:2509.19371cs.CLcs.AI2025-09EMNLP

提出知识注入的缩放定律,精准控制大模型知识融合量

How to inject knowledge efficiently? Knowledge Infusion Scaling Law for Pre-training Large Language Models

  • 通过小模型实验推导大模型最优知识注入量
  • 发现知识遗忘存在临界点,且随模型规模同步增长
  • 适合需要领域优化的大模型开发者参考

大型语言模型虽在多种下游任务中表现优异,但在特定领域基准测试中常因缺乏专业优化而表现不佳甚至产生幻觉。近期研究表明,在预训练阶段有策略地注入领域知识可显著提升性能。然而,如何平衡知识注入程度成为关键挑战:注入过少导致专业化不足,过多则引发灾难性遗忘。本文系统研究过度注入导致的记忆崩溃现象,发现两个关键规律:1)每个模型存在一个知识保留能力骤降的临界点;2)该临界点随模型规模一致增长。基于此,提出知识注入缩放定律,通过分析小模型特性预测大模型最优注入量。在不同模型规模和词元预算下的广泛实验验证了该定律的有效性与泛化性。

原文摘要 · Abstract (English)

Large language models (LLMs) have attracted significant attention due to their impressive general capabilities across diverse downstream tasks. However, without domain-specific optimization, they often underperform on specialized knowledge benchmarks and even produce hallucination. Recent studies show that strategically infusing domain knowledge during pretraining can substantially improve downstream performance. A critical challenge lies in balancing this infusion trade-off: injecting too little domain-specific data yields insufficient specialization, whereas excessive infusion triggers catastrophic forgetting of previously acquired knowledge. In this work, we focus on the phenomenon of memory collapse induced by over-infusion. Through systematic experiments, we make two key observations, i.e. 1) Critical collapse point: each model exhibits a threshold beyond which its knowledge retention capabilities sharply degrade. 2) Scale correlation: these collapse points scale consistently with the model's size. Building on these insights, we propose a knowledge infusion scaling law that predicts the optimal amount of domain knowledge to inject into large LLMs by analyzing their smaller counterparts. Extensive experiments across different model sizes and pertaining token budgets validate both the effectiveness and generalizability of our scaling law.

大模型知识注入缩放定律预训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。