arXiv:2602.23940cs.CLcs.LG2026-02中稿 · and presented at t…

对比多种BERT模型,为尼泊尔语句子级主题分类建立高效基准。

Benchmarking BERT-based Models for Sentence-level Topic Classification in Nepali Language

  • 测试10种BERT变体在尼泊尔语上的表现,聚焦多语言与印地语系模型。
  • MuRIL-large模型达90.60%的F1分数,显著优于其他模型。
  • 成果为尼泊尔语NLP研究提供可靠基线,适合低资源语言研究者参考。

基于Transformer的模型如BERT已在多种语言中显著推动自然语言处理发展。然而,使用天城文书写、资源较少的尼泊尔语仍相对未被充分研究。本研究对多语言、印地语系、印地语及尼泊尔语BERT变体进行基准测试,评估其在尼泊尔语主题分类中的有效性。十种预训练模型(包括mBERT、XLM-R、MuRIL、DevBERT、HindiBERT、IndicBERT和NepBERTa)在包含25,006个句子的平衡尼泊尔语数据集上进行微调与测试,涵盖五个概念领域,使用准确率、加权精确率、召回率、F1-score和AUROC进行评估。结果表明,印地语系模型(尤其是MuRIL-large)取得最高F1-score(90.60%),优于多语言与单语模型。NepBERTa也表现良好,F1-score达88.26%。这些发现为未来文档级分类及更广泛的尼泊尔语NLP应用建立了坚实基准。

原文摘要 · Abstract (English)

Transformer-based models such as BERT have significantly advanced Natural Language Processing (NLP) across many languages. However, Nepali, a low-resource language written in Devanagari script, remains relatively underexplored. This study benchmarks multilingual, Indic, Hindi, and Nepali BERT variants to evaluate their effectiveness in Nepali topic classification. Ten pre-trained models, including mBERT, XLM-R, MuRIL, DevBERT, HindiBERT, IndicBERT, and NepBERTa, were fine-tuned and tested on the balanced Nepali dataset containing 25,006 sentences across five conceptual domains and the performance was evaluated using accuracy, weighted precision, recall, F1-score, and AUROC metrics. The results reveal that Indic models, particularly MuRIL-large, achieved the highest F1-score of 90.60%, outperforming multilingual and monolingual models. NepBERTa also performed competitively with an F1-score of 88.26%. Overall, these findings establish a robust baseline for future document-level classification and broader Nepali NLP applications.

尼泊尔语BERT主题分类低资源语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。