arXiv:2512.21515cs.LGcs.CL2025-12ACL

用困惑度景观预测持续预训练效果,选对数据提升模型性能

Perplexity-Aware Data Scaling Law: Perplexity Landscapes Predict Performance for Continual Pre-training

  • 用预训练模型在领域数据上的困惑度,量化知识差距
  • 构建困惑度-性能关系模型,实现数据高效筛选
  • 适合医疗、通用领域模型优化,提升训练效率

持续预训练(CPT)是将基础模型适配到特定领域的重要方法。传统预训练缩放定律描述了数据规模与大语言模型测试损失之间的幂律关系,但单纯增加数据在CPT中边际收益迅速下降,导致数据利用不充分、训练效率低。为此,本文提出一种新的困惑度感知数据缩放定律,建立领域数据的困惑度景观与测试损失之间的预测关系。通过使用预训练模型在领域数据上的困惑度作为知识差距的代理指标,有效量化候选训练样本的信息困惑度景观。在不同困惑度区间拟合该缩放定律,实现高价值数据子集的自适应选择,优先保留能最大化知识吸收且减少冗余和噪声的内容。大量实验表明,该方法能持续识别接近最优的训练子集,在医疗和通用领域基准上均取得更优性能。

原文摘要 · Abstract (English)

Continual Pre-training (CPT) serves as a fundamental approach for adapting foundation models to domain-specific applications. Scaling laws for pre-training define a power-law relationship between dataset size and the test loss of an LLM. However, the marginal gains from simply increasing data for CPT diminish rapidly, yielding suboptimal data utilization and inefficient training. To address this challenge, we propose a novel perplexity-aware data scaling law to establish a predictive relationship between the perplexity landscape of domain-specific data and the test loss. Our approach leverages the perplexity derived from the pre-trained model on domain data as a proxy for estimating the knowledge gap, effectively quantifying the informational perplexity landscape of candidate training samples. By fitting this scaling law across diverse perplexity regimes, we enable adaptive selection of high-utility data subsets, prioritizing content that maximizes knowledge absorption while minimizing redundancy and noise. Extensive experiments demonstrate that our method consistently identifies near-optimal training subsets and achieves superior performance on both medical and general-domain benchmarks.

持续预训练数据筛选困惑度缩放定律

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。