arXiv:2410.05448cs.LGcs.CL2024-10被引 6

多任务学习可缩短语言模型的训练瓶颈期,提升学习效率。

Task Diversity Shortens the ICL Plateau

  • 同时训练多种多样ICL任务,打破学习平台期
  • 模型在多任务下学习速度加快,无须依赖单一任务复杂度
  • 适合研究大模型优化与自然语言多样性影响的学者

上下文学习(ICL)指语言模型根据输入示例和查询生成输出的能力。研究者常通过简化模型探究此能力,发现模型常经历长时间损失平台期,随后出现快速学习跃升。本文揭示:同时训练多个多样化的ICL任务可显著缩短该平台期,使每个任务更易学习。这一结果出人意料,因直觉上多任务应增加复杂度、延长学习过程。研究暗示,大模型的成功不仅源于海量数据,更得益于自然语言数据多样性带来的优化便利性。

原文摘要 · Abstract (English)

In-context learning (ICL) describes a language model's ability to generate outputs based on a set of input demonstrations and a subsequent query. To understand this remarkable capability, researchers have studied simplified, stylized models. These studies have consistently observed long loss plateaus, during which models exhibit minimal improvement, followed by a sudden, rapid surge of learning. In this work, we reveal that training on multiple diverse ICL tasks simultaneously shortens the loss plateaus, making each task easier to learn. This finding is surprising as it contradicts the natural intuition that the combined complexity of multiple ICL tasks would lengthen the learning process, not shorten it. Our result suggests that the recent success in large-scale training of language models may be attributed not only to the richness of the data at scale but also to the easier optimization (training) induced by the diversity of natural language training data.

上下文学习多任务学习模型训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。