通过优化知识分布,让大模型学得更广更深
Scaling LLM Knowledge Boundaries via Distribution-Optimized Synthesis

- 用知识密度引导合成数据,三阶段反馈实现精准生成
- 0.6B到16B模型均显示最优分布可最大化知识边界
- 适合想提升模型知识广度的研究者和工程师
通过合成数据注入知识对增强大语言模型至关重要。但现有方法仅以预设词元数量或固定数据比例为终止条件,缺乏对知识分布的认知,导致部分领域知识稀疏而其他领域冗余,限制了模型知识边界的扩展。本文从分布视角重新审视知识注入,提出假设:存在一种最优知识分布可最大化知识边界扩展。为此设计KDoS(知识分布优化合成)框架,引入知识密度概念,通过三阶段反馈机制,实现从盲目生成到分布优化的转变。构建基于维基百科、具有不同知识分布的合成数据集,在0.6B至16B参数量(Qwen、Ling、LLaMA)及1B至5B词元规模的数据上进行实验。关键发现:(1) 最优知识分布能持续最大化边界扩展;(2) 该分布跨模型架构与规模保持稳定;(3) KDoS在六项知识基准测试中均优于基线方法。本工作为合成数据驱动的知识注入提供了新视角与实用框架。
原文摘要 · Abstract (English)
Knowledge injection via synthetic data is crucial for enhancing Large Language Models (LLMs). However, current synthesis methods simply stop at preset token counts or fixed data ratios, lacking awareness of knowledge distribution. This results in some domains being sparse while others are redundant, limiting LLM knowledge boundaries. We revisit knowledge injection from a distribution perspective and hypothesize that an optimal knowledge distribution exists to maximize knowledge boundary expansion. We propose KDoS (Knowledge Distribution-optimized Synthesis), a framework that introduces knowledge density to drive synthesis through a three-stage feedback mechanism, shifting from blind generation to distribution-optimized synthesis. We construct Wikipedia-based synthetic data with varying knowledge distributions and conduct experiments on models from 0.6B to 16B (Qwen, Ling, LLaMA) and data scales from 1B to 5B tokens. Our key findings are: (1) an optimal knowledge distribution consistently maximizes boundary expansion; (2) this distribution is stable across backbones and scales; (3) KDoS outperforms baselines across six knowledge benchmarks. Our work offers a new perspective and practical framework for synthetic data-driven knowledge injection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。