用缩放定律动态优化预训练数据分布,省时省力提效果
Adaptive Data Optimization: Dynamic Sample Selection with Scaling Laws
- 基于各领域缩放定律在线评估数据潜力并调整混合比例
- 在不同计算规模下均实现媲美或超越已有方法的性能
- 无需额外模型或修改训练流程,适合实际部署
预训练数据组成是决定基础模型性能的关键因素,但目前缺乏在有限计算预算下分配资源的标准指南。现有方法多依赖小模型的大量实验或需代理模型进行动态调整,显著增加工作流复杂度和计算开销。本文提出自适应数据优化(ADO),一种在模型训练过程中在线优化数据分布的算法。与现有技术不同,ADO不依赖外部知识、代理模型或模型更新修改,而是利用各领域缩放定律估计训练中各领域的学习潜力,并据此动态调整数据混合比例,具备更强可扩展性和集成便利性。实验表明,ADO在不同计算规模下均能实现与先前方法相当或更优的性能,同时保持计算效率,为动态调整数据分布提供了实用方案。此外,ADO还通过缩放定律为数据采集策略提供了新视角。
原文摘要 · Abstract (English)
The composition of pretraining data is a key determinant of foundation models' performance, but there is no standard guideline for allocating a limited computational budget across different data sources. Most current approaches either rely on extensive experiments with smaller models or dynamic data adjustments that also require proxy models, both of which significantly increase the workflow complexity and computational overhead. In this paper, we introduce Adaptive Data Optimization (ADO), an algorithm that optimizes data distributions in an online fashion, concurrent with model training. Unlike existing techniques, ADO does not require external knowledge, proxy models, or modifications to the model update. Instead, ADO uses per-domain scaling laws to estimate the learning potential of each domain during training and adjusts the data mixture accordingly, making it more scalable and easier to integrate. Experiments demonstrate that ADO can achieve comparable or better performance than prior methods while maintaining computational efficiency across different computation scales, offering a practical solution for dynamically adjusting data distribution without sacrificing flexibility or increasing costs. Beyond its practical benefits, ADO also provides a new perspective on data collection strategies via scaling laws.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。