arXiv:2608.18094cs.CLcs.AI2026-08被引 2

为印度东北部9种低资源语言打造的多语言模型,显著提升这些语言的自然语言处理能力。

NE-BERT: A Multilingual Language Model for Nine Northeast Indian Languages

论文配图:NE-BERT: A Multilingual Language Model for Nine Northeast Indian Languages
图 1 · 摘自论文原文
  • 基于830万句语料训练,采用加权采样和自定义分词器提升性能
  • 在9种语言上比IndicBERT-V2和MuRIL平均困惑度降低15.97倍和7.64倍
  • 特别解决极低资源语言如Pnar、Kokborok的词汇碎片化问题,适合本地化NLP研究

大型预训练语言模型在多种语言中表现出色,但低资源语言仍被边缘化。本文提出NE-BERT,一个针对印度东北部9种语言及2种锚定语言(印地语、英语)的领域特定多语言编码模型,训练语料约830万句。通过加权数据采样和定制的SentencePiece Unigram分词器,NE-BERT在所有9种语言上均优于IndicBERT-V2和MuRIL,平均困惑度分别降低15.97倍和7.64倍,分词灵活性比mBERT高1.50倍。针对极端低资源语言如Pnar(1,002句)、Kokborok(2,463句),采用激进上采样策略缓解词汇碎片化问题。下游依存句法标注任务验证了其实际应用价值。模型、测试集与训练语料已按CC-BY-4.0协议开源,助力东北印度社区的NLP研究与数字包容。

原文摘要 · Abstract (English)

Large pretrained language models have demonstrated remarkable capabilities across diverse languages, yet critically underrepresented low-resource languages remain marginalized. We present NE-BERT, a domain-specific multilingual encoder model trained on approximately 8.3 million sentences spanning 9 Northeast Indian languages and 2 anchor languages (Hindi, English), a linguistically diverse region with minimal representation in existing multilingual models. By employing weighted data sampling and a custom SentencePiece Unigram tokenizer, NE-BERT outperforms IndicBERT-V2 and MuRIL across all 9 Northeast Indian languages, achieving 15.97X and 7.64X lower average perplexity respectively, with 1.50X better tokenization fertility than mBERT. We address critical vocabulary fragmentation issues in extremely low-resource languages such as Pnar (1,002 sentences) and Kokborok (2,463 sentences) through aggressive upsampling strategies. Downstream evaluation on part-of-speech tagging validates practical utility on three Northeast Indian languages. We release NE-BERT, test sets, and training corpus under CC-BY-4.0 to support NLP research and digital inclusion for Northeast Indian communities.

多语言模型低资源语言印度语言NLP开源

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。