首个支持僧伽罗语的开源大模型,让低资源语言也能用上先进AI。
SinLlama -- A Large Language Model for Sinhala
- 扩展Llama-3分词器并用1000万条僧伽罗语数据持续预训练
- 在文本分类任务中显著超越基础版Llama-3-8B模型
- 为僧伽罗语等低资源语言提供首个解码器架构开源模型
僧伽罗语等低资源语言常被开源大语言模型忽视。本研究将现有多语言模型Llama-3-8B扩展以更好服务僧伽罗语,通过添加僧伽罗语专属词汇增强分词器,并在清理后的1000万条僧伽罗语文本上进行持续预训练,形成Sinhala专用模型SinLlama。这是首个具备显式僧伽罗语支持的解码器架构开源大模型。当在三个文本分类任务上对SinLlama进行指令微调后,其性能显著优于Llama-3-8B的基础版与指令版。
原文摘要 · Abstract (English)
Low-resource languages such as Sinhala are often overlooked by open-source Large Language Models (LLMs). In this research, we extend an existing multilingual LLM (Llama-3-8B) to better serve Sinhala. We enhance the LLM tokenizer with Sinhala specific vocabulary and perform continual pre-training on a cleaned 10 million Sinhala corpus, resulting in the SinLlama model. This is the very first decoder-based open-source LLM with explicit Sinhala support. When SinLlama was instruction fine-tuned for three text classification tasks, it outperformed base and instruct variants of Llama-3-8B by a significant margin.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。