提出新方法评估真实与合成数据混合时的样本价值。
Data Value in the Age of Scaling: Understanding LLM Scaling Dynamics Under Real-Synthetic Data Mixtures
- 发现模型学习存在三阶段缩放规律,由两个断点界定。
- 理论推导出适用于混合数据的通用性边界,揭示关键影响因素。
- 方法高效准确,适合大规模数据集,计算开销小。
大型语言模型(LLM)的快速发展依赖于真实与合成数据的混合数据集。尽管合成数据具备可扩展性和低成本优势,但其生成机制(如top-p采样、温度调节和有限采样)常导致分布偏差,尤其在长尾知识上表现不足。本文揭示了模型学习中存在三个阶段的缩放行为,由两个断点定义,反映模型对头部与尾部知识的学习转变。我们推导出适用于真实-合成混合数据的LLM泛化界,揭示了影响泛化性能的关键因素。基于此理论,提出一种高效且可扩展的数据估值方法。在图像分类、情感分类、指令遵循和复杂推理四项任务上的实验表明,该方法在数据估值上优于现有最优基线,且计算成本显著降低。
原文摘要 · Abstract (English)
The rapid progress of large language models (LLMs) is fueled by the growing reliance on datasets that blend real and synthetic data. While synthetic data offers scalability and cost-efficiency, it often introduces systematic distributional discrepancies, particularly underrepresenting long-tail knowledge due to truncation effects from data generation mechanisms like top-p sampling, temperature scaling, and finite sampling. These discrepancies pose fundamental challenges in characterizing and evaluating the utility of mixed real-synthetic datasets. In this paper, we identify a three-phase scaling behavior characterized by two breakpoints that reflect transitions in model behavior across learning head and tail knowledge. We further derive an LLM generalization bound designed for real and synthetic mixtures, revealing several key factors that govern their generalization performance. Building on our theoretical findings, we propose an effective yet efficient data valuation method that scales to large-scale datasets. Comprehensive experiments across four tasks, including image classification, sentiment classification, instruction following, and complex reasoning, demonstrate that our method surpasses state-of-the-art baselines in data valuation with significantly low computational cost.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。