arXiv:2502.16802cs.CLcs.AI2025-02被引 5

按主题混合数据能让大模型预训练效果更好

Topic Over Source: The Key to Effective Data Mixing for Language Models Pre-training

  • 用多阶段方法自动标注数据主题,实现主题级数据混合
  • 不同混合策略下,主题混合均优于传统来源混合
  • 适合关注数据质量与预训练优化的研究者

大型语言模型的性能受预训练数据质量和构成的显著影响,这些数据涵盖多种语言、来源和主题。有效整合异构数据组对提升模型表现至关重要。以往研究多聚焦于基于来源的数据混合,忽视了数据的主题特征。为此,我们提出一种基于主题的数据混合策略,通过无监督聚类、LLM摘要生成与有监督分类器训练相结合的方式,自动获取细粒度主题标签。在此基础上,我们首次在多种混合策略(如RegMix、DoReMi、温度采样及基于下游任务的手动混合)中系统比较了主题与来源划分的效果。结果表明,按主题混合的数据使模型在各项指标上持续优于按来源混合的模型。理论分析显示,主题混合能实现更低的验证损失,提供更优的训练优化空间。相关代码、标注数据集及主题分类模型将公开发布,以支持后续研究。

原文摘要 · Abstract (English)

The performance of large language models (LLMs) is significantly affected by the quality and composition of their pre-training data, which is inherently diverse, spanning various languages, sources, and topics. Effectively integrating these heterogeneous data groups is crucial for optimizing LLM performance. Previous research has predominantly concentrated on source-based data mixing, often neglecting the nuanced topic-level characteristics of the data. To address this gap, we propose a topic-based data mixing strategy that utilizes detailed topic labels generated through a multi-stage process combining unsupervised clustering, LLM-based summarization, and supervised classifier training. With this strategy, we conduct the first comprehensive comparison of topic-based versus source-based partitioning across multiple mixing strategies. We demonstrate that language models pretrained on data mixed by topics consistently outperform those trained on data mixed by sources across multiple methods including RegMix, DoReMi,temperature-based sampling, and a manual mixing method based on downstream task performance. Our theoretical analysis reveals that topic-based data achieves significantly lower validation loss compared to source-based approaches, creating a better optimization landscape for model training. We will make our code, annotated datasets, and topic classification models publicly available to facilitate further research.

大模型预训练数据混合主题建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。