arXiv:2503.10995cs.CL2025-03ACL被引 21

TigerLLM是首个超越主流大模型的孟加拉语大模型家族。

TigerLLM - A Family of Bangla Large Language Models

  • 基于本土化数据构建孟加拉语大模型,提升语言适配性
  • 在标准评测中表现超越所有开源模型及GPT3.5
  • 为低资源语言建模提供可复现的新基准

大型语言模型的发展严重偏向英语及其他少数高资源语言。这一语言不平等在孟加拉语——全球第五大使用语言——中尤为突出。尽管已有少数尝试构建开源孟加拉语大模型,但其性能仍落后于高资源语言,且可复现性有限。为此,我们推出TigerLLM——一组孟加拉语大模型。实验表明,这些模型在标准基准测试中超越所有现有开源替代方案,并优于更大规模的专有模型如GPT3.5,确立了TigerLLM作为未来孟加拉语建模的新基准。

原文摘要 · Abstract (English)

The development of Large Language Models (LLMs) remains heavily skewed towards English and a few other high-resource languages. This linguistic disparity is particularly evident for Bangla - the 5th most spoken language. A few initiatives attempted to create open-source Bangla LLMs with performance still behind high-resource languages and limited reproducibility. To address this gap, we introduce TigerLLM - a family of Bangla LLMs. Our results demonstrate that these models surpass all open-source alternatives and also outperform larger proprietary models like GPT3.5 across standard benchmarks, establishing TigerLLM as the new baseline for future Bangla language modeling.

大模型低资源语言孟加拉语开源

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。