用低价流程为54种低资源语言训练通用小模型。
Kakugo: Distillation of Low-Resource Languages into Small Language Models
- 用大模型生成合成指令数据并翻译,构建训练集。
- 每语言总成本低于50美元,性能超越基础模型。
- 适合资源匮乏语言社区快速开发本地AI。
我们提出Kakugo,一种新颖且低成本的流程,仅需输入语言名称即可训练通用小型语言模型(SLMs)用于低资源语言。通过使用大型教师模型生成合成提示并翻译指令数据集,我们为54种低资源语言生成了训练数据和小型语言模型。在包括翻译、分类和问答在内的多样化自然语言处理任务中进行评估,结果表明该流程在各项任务中均持续优于基线模型。每种语言的生成与训练总成本低于50美元,为社区提供了一种可负担的语言特定AI开发方法。
原文摘要 · Abstract (English)
We present Kakugo, a novel and cost-effective pipeline designed to train general-purpose Small Language Models (SLMs) for low-resource languages using only the language name as input. By using a large teacher model to generate synthetic prompts and translate instruction datasets, we produced training data and SLMs for 54 low-resource languages. Evaluations across a diverse set of general natural language processing tasks, including translation, classification, and question answering, demonstrate that our pipeline consistently improves performance over base models. With a total generation and training cost of under $50 per language, Kakugo offers an accessible method for communities to develop language-specific AI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。