arXiv:2501.18962cs.LG2025-01NeurIPS被引 3

优化合成数据迭代训练的预算分配,提升模型性能上限

Spend Wisely: Maximizing Post-Training Gains in Iterative Synthetic Data Bootstrapping

  • 提出增长型预算分配策略,优于固定预算
  • 指数增长策略在图像去噪和数学推理任务中表现更优
  • 适合追求高效迭代训练的模型开发者

现代基础模型常在后训练阶段采用迭代式自举:模型生成合成数据,外部验证器过滤低质量样本,高质量子集用于后续微调。多轮迭代后模型性能提升,关键问题是:如何在各轮次间分配生成与训练总预算以最大化最终性能?本文建立理论框架分析预算分配策略。结果表明,固定预算策略高概率无法收敛,而增长型策略——尤其是指数增长——具有显著理论优势。在扩散概率模型图像去噪与大语言模型数学推理任务上的实验显示,指数与多项式增长策略均持续优于固定策略,其中指数策略表现更稳定。

原文摘要 · Abstract (English)

Modern foundation models often undergo iterative ``bootstrapping'' in their post-training phase: a model generates synthetic data, an external verifier filters out low-quality samples, and the high-quality subset is used for further fine-tuning. Over multiple iterations, the model performance improves, raising a crucial question: How should the total budget for generation and training be allocated across iterations to maximize final performance? In this work, we develop a theoretical framework for analyzing budget allocation strategies. Specifically, we show that constant policies fail to converge with high probability, while increasing policies -- particularly exponential growth policies -- exhibit significant theoretical advantages. Experiments on image denoising with diffusion probabilistic models and math reasoning with large language models show that both exponential and polynomial growth policies consistently outperform constant policies, with exponential policies often providing more stable performance.

模型训练合成数据预算分配迭代优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。