用元学习自动评估数据价值,提升训练效率
DataRater: Meta-Learned Dataset Curation
- 通过元梯度学习预测每个数据点的训练价值
- 在多种模型和数据集上显著提升计算效率
- 适合需要高效数据筛选的研究者与工程师
基础模型的性能高度依赖于训练数据质量。现有数据筛选方法多依赖人工调整粗粒度数据组合或手工设计规则,难以扩展。本文提出DataRater,一种基于元学习的数据价值评估方法,通过元梯度优化,在未见数据上提升训练效率。该方法可自动衡量任一数据点的训练价值。在多种模型规模与数据集上的实验表明,使用DataRater进行数据过滤能显著提高计算效率。
原文摘要 · Abstract (English)
The quality of foundation models depends heavily on their training data. Consequently, great efforts have been put into dataset curation. Yet most approaches rely on manual tuning of coarse-grained mixtures of large buckets of data, or filtering by hand-crafted heuristics. An approach that is ultimately more scalable (let alone more satisfying) is to \emph{learn} which data is actually valuable for training. This type of meta-learning could allow more sophisticated, fine-grained, and effective curation. Our proposed \emph{DataRater} is an instance of this idea. It estimates the value of training on any particular data point. This is done by meta-learning using `meta-gradients', with the objective of improving training efficiency on held out data. In extensive experiments across a range of model scales and datasets, we find that using our DataRater to filter data is highly effective, resulting in significantly improved compute efficiency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。