arXiv:2410.08942cs.LGcs.AI2024-10ICLR被引 6

用随机矩阵理论分析合成数据优劣,揭示何时能提升模型性能。

Maximizing the Potential of Synthetic Data: Insights from Random Matrix Theory

  • 基于随机矩阵理论分析混合数据下的分类器表现
  • 发现合成数据提升性能需满足生成质量与验证策略条件
  • 揭示标签噪声存在平滑相变,适合关注数据质量的开发者

合成数据在训练大语言模型中备受关注,但低质量数据会损害模型性能(如Shumailov等,2023;Seddik等,2024)。一种潜在解决方案是数据剪枝,即根据评分函数(人工或机器反馈)保留高质量数据。先前工作Feng等(2024)研究了随样本量增加时合成数据训练模型的表现。本文通过随机矩阵理论,推导了在高维设定下,由真实数据与剪枝后合成数据混合训练的二分类器性能。研究揭示了合成数据可提升性能的条件,关键在于生成模型的质量与验证策略。同时,我们发现合成标签噪声存在平滑相变,与以往无穷样本极限下的突变行为不同。小规模模型和大语言模型的实验验证了理论结果。

原文摘要 · Abstract (English)

Synthetic data has gained attention for training large language models, but poor-quality data can harm performance (see, e.g., Shumailov et al. (2023); Seddik et al. (2024)). A potential solution is data pruning, which retains only high-quality data based on a score function (human or machine feedback). Previous work Feng et al. (2024) analyzed models trained on synthetic data as sample size increases. We extend this by using random matrix theory to derive the performance of a binary classifier trained on a mix of real and pruned synthetic data in a high dimensional setting. Our findings identify conditions where synthetic data could improve performance, focusing on the quality of the generative model and verification strategy. We also show a smooth phase transition in synthetic label noise, contrasting with prior sharp behavior in infinite sample limits. Experiments with toy models and large language models validate our theoretical results.

合成数据随机矩阵模型性能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。