用多维度评分融合提升预训练数据质量,少用数据却效果更好
FIRE: Flexible Integration of Data Quality Ratings for Effective Pre-Training
- 融合多种数据质量评分,统一为综合信号
- 仅需随机采样37.5%数据量即达目标性能
- 适合追求高效预训练的模型开发者
选择高质量数据可提升大语言模型预训练效率。现有方法多依赖启发式策略或单一质量信号,难以全面评估数据质量。本文提出FIRE框架,灵活整合多种数据质量评估器,将不同质量信号对齐至统一空间,为每个数据点生成综合质量评分。进一步设计基于FIRE的渐进式数据选择方案,迭代优化高质量数据选取。大量实验表明,FIRE优于其他数据选择方法,在多个下游任务上显著提升预训练模型性能,且所需数据量不足随机基线的37.5%即可达到目标性能。
原文摘要 · Abstract (English)
Selecting high-quality data can improve the pretraining efficiency of large language models (LLMs). Existing methods generally rely on heuristic techniques or single quality signals, limiting their ability to evaluate data quality comprehensively. In this work, we propose FIRE, a flexible and scalable framework for integrating multiple data quality raters, which allows for a comprehensive assessment of data quality across various dimensions. FIRE aligns multiple quality signals into a unified space, and integrates diverse data quality raters to provide a comprehensive quality signal for each data point. Further, we introduce a progressive data selection scheme based on FIRE that iteratively refines the selection of high-quality data points. Extensive experiments show that FIRE outperforms other data selection methods and significantly boosts pretrained model performance across a wide range of downstream tasks, while requiring less than 37.5\% of the training data needed by the Random baseline to reach the target performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。