arXiv:2410.08102cs.CL2024-10ACL被引 12

多角色协作筛选训练数据,提速语言模型预训练。

Efficient Pretraining Data Selection for Language Models via Multi-Actor Collaboration

  • 多个数据筛选方法自主评分并动态调整规则,像独立演员协同工作。
  • 在多个基准上平均性能提升达10.5%,收敛速度明显加快。
  • 适合追求高效预训练的语言模型研究者和工业应用开发者。

高效的训练数据选择对加速语言模型(LM)预训练至关重要。尽管已有多种方法提升数据效率,但针对不同方法间内在冲突的研究仍有限,难以实现最优数据选择。为此,我们提出一种多角色协作的数据选择机制:每个数据选择方法基于自身标准独立为数据排序,并利用当前模型状态更新其排序规则,作为独立的数据选择‘演员’;同时设计一个控制台,在预训练的不同阶段动态调节各‘演员’的影响权重,并持续整合所有演员的信息。我们进行了广泛的实验评估该框架。结果表明,该方法显著提升了数据效率,加速了语言模型预训练的收敛速度,在多个语言模型基准上相比现有最佳方法实现了最高达10.5%的平均相对性能提升。

原文摘要 · Abstract (English)

Efficient data selection is crucial to accelerate the pretraining of language model (LMs). While various methods have been proposed to enhance data efficiency, limited research has addressed the inherent conflicts between these approaches to achieve optimal data selection for LM pretraining. To tackle this problem, we propose a multi-actor collaborative data selection mechanism: each data selection method independently prioritizes data based on its criterion and updates its prioritization rules using the current state of the model, functioning as an independent actor for data selection; and a console is designed to adjust the impacts of different actors at various stages and dynamically integrate information from all actors throughout the LM pretraining process. We conduct extensive empirical studies to evaluate our multi-actor framework. The experimental results demonstrate that our approach significantly improves data efficiency, accelerates convergence in LM pretraining, and achieves an average relative performance gain up to $10.5\%$ across multiple language model benchmarks compared to the state-of-the-art methods.

语言模型数据筛选预训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。