arXiv:2507.16178cs.LGcs.AI2025-07ICML被引 6

动态调整数据权重,让大模型训练更高效

LLM Data Selection and Utilization via Dynamic Bi-level Optimization

  • 用双层优化动态调整每批数据权重,适应模型变化
  • 在随机选数据下提升模型性能,且可迁移至其他方法
  • 揭示模型训练中数据偏好演化规律,适合优化训练者

大规模训练数据对开发强大语言模型至关重要,但如何选择高质量数据已成为提升训练效率、降低计算成本的关键。现有数据选择方法多依赖静态、与训练无关的标准,未能捕捉模型训练过程中数据的动态交互。本文提出一种数据加权模型(DWM),在每批次训练中动态调整所选数据的权重,实现动态数据利用。特别地,采用双层优化框架来更新加权模型,以更好捕捉模型在训练过程中的动态数据偏好。实验表明,DWM能显著提升使用随机选取数据训练的模型性能,且学习到的加权模型可迁移用于增强其他数据选择方法及不同规模的模型。此外,我们进一步分析了模型数据偏好随训练的演变过程,为理解训练中数据偏好提供了新视角。

原文摘要 · Abstract (English)

While large-scale training data is fundamental for developing capable large language models (LLMs), strategically selecting high-quality data has emerged as a critical approach to enhance training efficiency and reduce computational costs. Current data selection methodologies predominantly rely on static, training-agnostic criteria, failing to account for the dynamic model training and data interactions. In this paper, we propose a new Data Weighting Model (DWM) to adjust the weight of selected data within each batch to achieve a dynamic data utilization during LLM training. Specially, to better capture the dynamic data preference of the trained model, a bi-level optimization framework is implemented to update the weighting model. Our experiments demonstrate that DWM enhances the performance of models trained with randomly-selected data, and the learned weighting model can be transferred to enhance other data selection methods and models of different sizes. Moreover, we further analyze how a model's data preferences evolve throughout training, providing new insights into the data preference of the model during training.

数据选择大模型训练动态优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。