为孟加拉语构建首个专用大模型,填补低资源语言空白
BanglaLlama: LLaMA for Bangla Language
- 构建22.4万条高质量孟加拉语指令数据集
- 推出5个基线与指令微调模型,显著提升处理能力
- 适合关注低资源语言与多语言AI的研究者
孟加拉语是全球第五大使用语言,拥有约2.4亿母语者和3亿使用者。尽管如此,该语言仍属“低资源”范畴,现有预训练模型在孟加拉语处理任务中表现不佳。本文通过:(1) 构建两个高质量的孟加拉语指令数据集,共计22.4万条样本——Bangla-Orca(17.2万条)和Bangla-Alpaca(5.2万条);(2) 利用这些数据集训练出一系列开源的孟加拉语专用大模型,即BanglaLlama,包含五个基础模型与指令微调版本。本文详述方法、发布两个大规模数据集,并在多个基准上展示模型性能。我们相信,本研究提供的数据集与模型将为未来孟加拉语研究树立新标准。
原文摘要 · Abstract (English)
Bangla is a language spoken by approximately 240 million native speakers and around 300 million people worldwide. Despite being the 5th largest spoken language in the world, Bangla is still a "low-resource" language, and existing pretrained language models often struggle to perform well on Bangla Language Processing (BLP) tasks. This paper addresses this gap by: (1) introducing two high-quality translated Bangla-instruction datasets totaling 224k samples - Bangla-Orca (172k) and Bangla-Alpaca (52k); and (2) leveraging these datasets to develop BanglaLlama, an open-source family of Bangla-specific LLMs, consisting of five base and instruct variants. We present our methodology, two large datasets, and comprehensive benchmarking results showcasing the effectiveness of our dataset and model on multiple benchmarks. We believe our proposed datasets and models will serve as the new standard baseline for future research focused on this widely spoken yet "low-resource" language.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。