arXiv:2509.09101cs.CL2025-09被引 17

首个专为孟加拉语代码生成设计的LLM系列,提升低资源语言编程能力。

TigerCoder: A Novel Suite of LLMs for Code Generation in Bangla

  • 构建首个孟加拉语代码指令数据集,支持编程领域适配
  • 推出MBPP-Bangla基准,实现11%-18%性能提升
  • 开源全套资源,助力低资源语言模型研究

尽管孟加拉语是全球第五大使用语言,但在大型语言模型(LLMs)中仍严重缺乏代码生成支持,主要受限于高质量训练数据稀缺。为此,我们首次提出专用于孟加拉语的代码生成大模型家族(TigerCoder,含1B与9B参数版本)。贡献包括:(1) 构建全面的孟加拉语代码指令数据集,用于编程领域适配;(2) 发布MBPP-Bangla评估基准,衡量孟加拉语代码生成能力;(3) TigerCoder系列模型在Pass@1指标上相比现有多语言及通用孟加拉语LLM显著提升11%-18%。实验表明,精心构建的高质量数据可有效弥补小模型在低资源语言上的局限。所有资源均已开源,以推动孟加拉语LLM研究发展。

原文摘要 · Abstract (English)

Despite being the 5th most spoken language, Bangla remains underrepresented in Large Language Models (LLMs), particularly for code generation. This primarily stems from the scarcity of high-quality data to pre-train and/or finetune such models. Hence, we introduce the first dedicated family of Code LLMs for Bangla (1B & 9B). We offer three major contributions: (1) a comprehensive Bangla code instruction datasets for programming domain adaptation; (2) MBPP-Bangla, an evaluation benchmark for Bangla code generation; and (3) the TigerCoder-family of Code LLMs, achieving significant ~11-18% performance gains at Pass@1 over existing multilingual and general-purpose Bangla LLMs. Our findings show that curated, high-quality datasets can overcome limitations of smaller models for low-resource languages. We open-source all resources to advance further Bangla LLM research.

代码生成低资源语言大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。