arXiv:2505.14070cs.CL2025-05AAAI被引 3

用知识密度筛选数据,提升大模型的知识能力

Enhancing LLMs via High-Knowledge Data Selection

  • 设计无梯度的知识评分器,从知识丰富度选数据
  • 在多领域知识库中评估文本知识密度与覆盖度
  • 适用于通用和特定领域的知识增强,效果显著

大语言模型的性能与其训练数据质量密切相关。尽管已有研究提出高质量数据筛选方法,但均未充分考虑语料中知识丰富度的重要性。本文提出一种全新的无梯度高知识评分器(HKS),从知识维度筛选高质量数据,以缓解预训练语料中的知识匮乏问题。我们构建了跨领域的综合知识元素池,并引入知识密度和知识覆盖度作为衡量文本知识含量的指标。基于此,设计综合知识评分器,可筛选出知识密集型数据,亦可通过限定知识元素范围实现特定领域高知识数据的选择。在高知识双语数据集上训练模型,实验表明该评分器显著提升了模型在知识密集型及通用理解任务上的表现,有效增强了模型的通用与领域专用能力。

原文摘要 · Abstract (English)

The performance of Large Language Models (LLMs) is intrinsically linked to the quality of its training data. Although several studies have proposed methods for high-quality data selection, they do not consider the importance of knowledge richness in text corpora. In this paper, we propose a novel and gradient-free High-Knowledge Scorer (HKS) to select high-quality data from the dimension of knowledge, to alleviate the problem of knowledge scarcity in the pre-trained corpus. We propose a comprehensive multi-domain knowledge element pool and introduce knowledge density and coverage as metrics to assess the knowledge content of the text. Based on this, we propose a comprehensive knowledge scorer to select data with intensive knowledge, which can also be utilized for domain-specific high-knowledge data selection by restricting knowledge elements to the specific domain. We train models on a high-knowledge bilingual dataset, and experimental results demonstrate that our scorer improves the model's performance in knowledge-intensive and general comprehension tasks, and is effective in enhancing both the generic and domain-specific capabilities of the model.

大模型数据筛选知识增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。