arXiv:2412.06738cs.CL2024-12中稿 · PACLIC38

用大模型生成日语数据,让小模型少样本学得更好

JAPAGEN: Efficient Few/Zero-shot Learning via Japanese Training Dataset Generation with LLM

  • 用大模型合成日语训练数据,支持少样本和零样本学习
  • 在6个日语任务上,小模型性能媲美传统提示方法
  • 适合资源有限但需日语自然语言处理的场景

近期一些研究指出大型语言模型(LLMs)可作为高效的监督训练数据生成器,具有提升推理效率和降低数据收集成本的优势。然而,这些研究主要集中在英语任务上。本文回答一个基础问题:LLMs能否成为其他语言任务的有效训练数据生成器?我们利用LLMs在六个不同的日语下游任务中,于少样本和零样本学习场景下合成监督训练数据,并用这些数据训练紧凑模型(如BERT)。该新方法称为JAPAGEN。实验结果表明,JAPAGEN在需要正式文本输入的分类任务中表现稳健,性能与传统LLM提示策略相当。

原文摘要 · Abstract (English)

Recently some studies have highlighted the potential of Large Language Models (LLMs) as effective generators of supervised training data, offering advantages such as enhanced inference efficiency and reduced costs associated with data collection. However, these studies have predominantly focused on English language tasks. In this paper, we address the fundamental research question: Can LLMs serve as proficient training data generators for other language tasks? Specifically, we leverage LLMs to synthesize supervised training data under few-shot and zero-shot learning scenarios across six diverse Japanese downstream tasks. Subsequently, we utilize this synthesized data to train compact models (e.g., BERT). This novel methodology is termed JAPAGEN. Our experimental findings underscore that JAPAGEN achieves robust performance in classification tasks that necessitate formal text inputs, demonstrating competitive results compared to conventional LLM prompting strategies.

少样本学习日语NLP数据生成高效训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。