用大模型生成数据,让小模型在低资源语言中表现超越大模型。
Better as Generators Than Classifiers: Leveraging LLMs and Synthetic Data for Low-Resource Multilingual Classification
- 用多语言大模型生成11种语言的合成数据,训练小型模型。
- 少量合成数据就使小模型在低资源语言上超越大模型。
- 适合需要轻量高效多语言系统的开发者和研究者。
大型语言模型(LLMs)展现出卓越的多语言能力,使其在高资源与低资源语言中均具应用潜力。一个关键应用场景是生成合成样本,用于在人工标注数据稀缺的低资源场景中训练小型模型。本文探究这些合成数据生成能力能否作为知识蒸馏手段,生成性能媲美甚至优于大型LLM的小型模型。我们使用最先进的多语言LLM生成覆盖11种语言和4类分类任务的合成数据集,再通过微调或指令调优训练小型模型,或将合成数据作为紧凑型LLM的上下文示例。实验表明,即使少量合成数据也能使小型模型在低资源语言上超越生成器本身。整体结果表明,应将LLM主要视为生成器(教师),而非分类器,其生成的数据可赋能更小、更高效的多语言模型。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have demonstrated remarkable multilingual capabilities, making them promising tools in both high- and low-resource languages. One particularly valuable use case is generating synthetic samples that can be used to train smaller models in low-resource scenarios where human-labelled data is scarce. In this work, we investigate whether these synthetic data generation capabilities can serve as a form of distillation, producing smaller models that perform on par with or even better than massive LLMs across languages and tasks. To this end, we use a state-of-the-art multilingual LLM to generate synthetic datasets covering 11 languages and 4 classification tasks. These datasets are then used to train smaller models via fine-tuning or instruction tuning, or as synthetic in-context examples for compact LLMs. Our experiments show that even small amounts of synthetic data enable smaller models to outperform the large generator itself, particularly in low-resource languages. Overall, the results suggest that LLMs are best utilised as generators (teachers) rather than classifiers, producing data that empowers smaller and more efficient multilingual models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。