arXiv:2502.11187cs.CLcs.AI2025-02ACL被引 16

首个10亿与30亿参数的孟加拉语大模型,填补低资源语言空白

TituLLMs: A Family of Bangla LLMs with Comprehensive Benchmarking

  • 基于370亿词元训练,扩展Llama分词器适配孟加拉语文化特征
  • 自建5个基准数据集,证明其在孟加拉语任务上优于多语言基线模型
  • 开源模型与数据集,助力低资源语言大模型研究

本文提出TituLLMs,首个针对孟加拉语的大规模预训练语言模型,包含10亿和30亿参数版本。受训练与推理计算限制,我们聚焦于小型模型。为训练,我们构建了约370亿词元的预训练数据集,并扩展Llama-3.2分词器以融入语言与文化特异性知识,提升训练与推理效率。由于缺乏孟加拉语评测数据集,我们开发了五个基准测试集。对多种模型(包括TituLLMs)进行评测,结果表明其在多数任务上超越初始多语言版本,但并非总是如此,凸显语言适配的复杂性。本工作为将现有多语言开源模型适配至其他低资源语言奠定基础。为促进广泛采用与后续研究,我们已公开发布TituLLMs模型及评测数据集(https://huggingface.co/collections/hishab/titulm-llama-family-6718d31fc1b83529276f490a)。

原文摘要 · Abstract (English)

In this paper, we present TituLLMs, the first large pretrained Bangla LLMs, available in 1b and 3b parameter sizes. Due to computational constraints during both training and inference, we focused on smaller models. To train TituLLMs, we collected a pretraining dataset of approximately ~37 billion tokens. We extended the Llama-3.2 tokenizer to incorporate language- and culture-specific knowledge, which also enables faster training and inference. There was a lack of benchmarking datasets to benchmark LLMs for Bangla. To address this gap, we developed five benchmarking datasets. We benchmarked various LLMs, including TituLLMs, and demonstrated that TituLLMs outperforms its initial multilingual versions. However, this is not always the case, highlighting the complexities of language adaptation. Our work lays the groundwork for adapting existing multilingual open models to other low-resource languages. To facilitate broader adoption and further research, we have made the TituLLMs models and benchmarking datasets publicly available (https://huggingface.co/collections/hishab/titulm-llama-family-6718d31fc1b83529276f490a).

孟加拉语大模型低资源语言开源

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。