arXiv:2505.17102cs.CL2025-05EMNLP被引 6

首个针对孟加拉语的字节级模型,提升形态丰富语言处理效果

BanglaByT5: Byte-Level Modelling for Bangla

  • 采用字节级建模,避免传统分词器对孟加拉语的碎片化处理
  • 在14GB高质量语料上预训练,零样本和有监督任务均表现优异
  • 轻量高效,适合资源受限与可扩展场景下的孟加拉语NLP应用

大型语言模型在自然语言处理任务中取得显著成功,但多数使用BPE、SentencePiece等传统分词器,难以捕捉孟加拉语这类形态丰富的语言细节。本文提出BanglaByT5,首个专为孟加拉语设计的字节级编码器-解码器模型,基于谷歌ByT5的小型变体,在14GB精心筛选的文学与新闻文章语料上进行预训练。通过生成与分类任务的零样本及有监督评估,该模型表现竞争力,超越多个多语言及更大规模模型。结果表明,字节级建模对形态丰富的语言有效,且BanglaByT5具备轻量高效潜力,适用于资源受限与可扩展环境中的孟加拉语自然语言处理。

原文摘要 · Abstract (English)

Large language models (LLMs) have achieved remarkable success across various natural language processing tasks. However, most LLM models use traditional tokenizers like BPE and SentencePiece, which fail to capture the finer nuances of a morphologically rich language like Bangla (Bengali). In this work, we introduce BanglaByT5, the first byte-level encoder-decoder model explicitly tailored for Bangla. Built upon a small variant of Googles ByT5 architecture, BanglaByT5 is pre-trained on a 14GB curated corpus combining high-quality literary and newspaper articles. Through zeroshot and supervised evaluations across generative and classification tasks, BanglaByT5 demonstrates competitive performance, surpassing several multilingual and larger models. Our findings highlight the efficacy of byte-level modelling for morphologically rich languages and highlight BanglaByT5 potential as a lightweight yet powerful tool for Bangla NLP, particularly in both resource-constrained and scalable environments.

孟加拉语字节级建模轻量模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。