arXiv:2603.19149cs.CLcs.LG2026-03被引 2

通过计算规模定律优化模型拆分策略,提升多领域语言模型性能。

Optimal Splitting of Language Models from Mixtures to Specialized Domains

  • 独立预训练多个模型,用规模定律确定最优算力分配。
  • 准确预测不同参数量与训练数据下的模型损失,支持扩展至更大规模。
  • 适用于不同算力预算和模型尺寸,显著提升常识推理表现。

语言模型因预训练数据的规模和多样性,在各类知识、语言和推理任务上表现优异。标准训练流程为两阶段:先在全量数据上预训练,再在高质量细分数据上微调。在多领域场景中,需对多个模型分别在各专用领域持续预训练,称为拆分模型训练。本文提出一种方法:在通用预训练语料上独立预训练多个模型,并利用规模定律确定预训练与持续预训练之间的最优算力分配。该方法可准确预测参数量为 N、预训练数据量为 D、持续预训练数据量为 D' 的模型损失,并外推至更大模型尺寸和更大数据量。应用于语言模型训练时,该方法在不同模型尺寸和算力预算下,均一致提升常识知识与推理基准的表现。

原文摘要 · Abstract (English)

Language models achieve impressive performance on a variety of knowledge, language, and reasoning tasks due to the scale and diversity of pretraining data available. The standard training recipe is a two-stage paradigm: pretraining first on the full corpus of data followed by specialization on a subset of high quality, specialized data from the full corpus. In the multi-domain setting, this involves continued pretraining of multiple models on each specialized domain, referred to as split model training. We propose a method for pretraining multiple models independently over a general pretraining corpus, and determining the optimal compute allocation between pretraining and continued pretraining using scaling laws. Our approach accurately predicts the loss of a model of size N with D pretraining and D' specialization tokens, and extrapolates to larger model sizes and number of tokens. Applied to language model training, our approach improves performance consistently across common sense knowledge and reasoning benchmarks across different model sizes and compute budgets.

模型拆分规模定律语言模型算力优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。