arXiv:2503.01506cs.CL2025-03EMNLP被引 9

通过样本级质量与多样性评估,动态优化预训练数据混合策略。

SampleMix: A Sample-wise Pre-training Data Mixing Strategey by Coordinating Data Quality and Diversity

  • 自下而上按样本质量与多样性全局采样,打破领域固定权重限制。
  • 在多任务下游评测中超越传统方法,但需1.4至2.1倍训练步数达基线性能。
  • 适合追求高质量预训练数据分布的模型开发者,尤其在数据混合适配场景。

现有大语言模型预训练数据混合方法多采用领域级策略,先设定领域权重,再在各领域内均匀采样。此类方法忽略领域间重叠与共性,难以控制全局数据多样性;且领域内均匀采样忽视样本级特征,导致数据分布欠优。为此,本文提出一种基于自下而上范式的样本级数据混合方法。该方法通过系统评估每个样本的质量与多样性,实现跨领域的全局采样,动态确定最优领域分布。在多个下游任务及困惑度评估中,SampleMix表现优于现有领域级方法。同时,其达到基线性能所需训练步数为1.4至2.1倍,凸显了在预训练数据优化方面的巨大潜力。

原文摘要 · Abstract (English)

Existing pretraining data mixing methods for large language models (LLMs) typically follow a domain-wise methodology, a top-down process that first determines domain weights and then performs uniform data sampling across each domain. However, these approaches neglect significant inter-domain overlaps and commonalities, failing to control the global diversity of the constructed training dataset. Further, uniform sampling within domains ignores fine-grained sample-specific features, potentially leading to suboptimal data distribution. To address these shortcomings, we propose a novel sample-wise data mixture approach based on a bottom-up paradigm. This method performs global cross-domain sampling by systematically evaluating the quality and diversity of each sample, thereby dynamically determining the optimal domain distribution. Comprehensive experiments across multiple downstream tasks and perplexity assessments demonstrate that SampleMix surpasses existing domain-based methods. Meanwhile, SampleMix requires 1.4x to 2.1x training steps to achieves the baselines' performance, highlighting the substantial potential of SampleMix to optimize pre-training data.

预训练数据混合语言模型样本选择

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。