arXiv:2607.23344cs.CLcs.LG2026-07

针对马拉地语的低资源命名实体识别,微调的BERT模型优于通用大模型。

BERT-based Models vs. Large Language Models for Low-Resource Named Entity Recognition: A Comparative Study on Marathi

  • 用马拉地语数据微调MahaBERT-v2模型,针对性提升识别效果。
  • 微调模型F1最高达0.91,显著超越通用大模型的0.57~0.69。
  • 适合低资源语言研究者参考,强调专用模型的重要性。

马拉地语等低资源语言的命名实体识别(NER)因标注数据有限和语言复杂性而面临挑战。尽管大型语言模型(LLMs)在多种自然语言处理任务中表现优异,但在低资源场景下的语言特定NER效果尚不明确。本研究在MahaNER数据集的不同变体上微调MahaBERT-v2,并与现有基准模型及主流通用大模型(Gemini、LLaMA-3.3-70B、Gemma)进行系统对比。所有模型在马拉地语NER测试集上以精确率、召回率和F1-score评估。实验结果表明,微调后的MahaBERT模型始终优于基线和所有评估的LLMs,F1分数范围为0.88至0.91,超过现有MahaNER模型(0.8843),显著高于大模型方法(F1为0.57至0.69)。研究证明,在低资源语言处理中,基于领域相关数据训练的任务特定语言模型仍比通用大模型更有效,凸显专用架构的持续价值。

原文摘要 · Abstract (English)

Named Entity Recognition (NER) for low-resource languages such as Marathi remains a challenging task due to limited annotated resources and linguistic complexity. Although recent Large Language Models (LLMs) have demonstrated strong performance across a wide range of natural language processing tasks, their effectiveness for language-specific NER in low-resource settings remains uncertain. In this study, we fine-tune MahaBERT-v2 on different variants of the MahaNER dataset and systematically compare the performance of these models with an existing MahaNER baseline and prominent general-purpose LLMs, including Gemini, LLaMA-3.3-70B, and Gemma models. All models are evaluated on a Marathi NER test dataset using standard metrics of precision, recall, and F1-score. The experimental results show that the fine-tuned MahaBERT-based models consistently outperform both the baseline and all evaluated LLMs, with the fine-tuned models achieving F1-scores ranging from 0.88 to 0.91, surpassing the existing MahaNER model (0.8843) and significantly exceeding the performance of LLM-based approaches, whose F1-scores range from 0.57 to 0.69. These findings demonstrate that task-specific, language-focused models trained on domain-relevant data remain more effective than general-purpose LLMs for Marathi NER, highlighting the continued importance of specialized architectures for low-resource language processing.

命名实体识别低资源语言BERT大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。