arXiv:2502.11198cs.CLcs.LG2025-02被引 5

首个孟加拉方言命名实体识别基准数据集,助力低资源语言研究

ANCHOLIK-NER: A Benchmark Dataset for Bangla Regional Named Entity Recognition

  • 构建首个孟加拉方言NER基准数据集,覆盖5个地区共1.7万句
  • 多语言BERT在莫因辛格方言上达F1 82.611%最佳表现
  • 揭示查特格龙等方言识别难题,为方言模型优化提供方向

区域方言的命名实体识别(NER)是自然语言处理中的关键但研究不足领域,尤其对孟加拉语等低资源语言而言。尽管标准孟加拉语的NER系统已有进展,但现有资源与模型均未针对巴里沙尔、吉大港、莫因辛格、诺阿哈利和锡尔赫特等方言开展工作,这些方言具有独特语言特征,现有模型难以有效处理。为此,我们提出ANCHOLIK-NER,首个面向孟加拉方言的命名实体识别基准数据集,包含17,405条句子,覆盖五个地区。数据源自公开资源并经人工翻译校准,确保实体一致性。我们在该数据集上评估了三种基于Transformer的模型:Bangla BERT、Bangla BERT Base 和 BERT Base Multilingual Cased。结果显示,BERT Base Multilingual Cased 在跨区域识别中表现最优,在莫因辛格方言上取得82.611%的F1分数。尽管整体表现良好,吉大港方言仍存在精度与召回率偏低问题。由于此前无专门针对孟加拉方言的NER系统,本工作为填补空白奠定基础。未来将聚焦提升薄弱地区的性能,并扩展更多方言,推动方言感知型NER系统的发展。

原文摘要 · Abstract (English)

Named Entity Recognition (NER) in regional dialects is a critical yet underexplored area in Natural Language Processing (NLP), especially for low-resource languages like Bangla. While NER systems for Standard Bangla have made progress, no existing resources or models specifically address the challenge of regional dialects such as Barishal, Chittagong, Mymensingh, Noakhali, and Sylhet, which exhibit unique linguistic features that existing models fail to handle effectively. To fill this gap, we introduce ANCHOLIK-NER, the first benchmark dataset for NER in Bangla regional dialects, comprising 17,405 sentences distributed across five regions. The dataset was sourced from publicly available resources and supplemented with manual translations, ensuring alignment of named entities across dialects. We evaluate three transformer-based models - Bangla BERT, Bangla BERT Base, and BERT Base Multilingual Cased - on this dataset. Our findings demonstrate that BERT Base Multilingual Cased performs best in recognizing named entities across regions, with significant performance observed in Mymensingh with an F1-score of 82.611%. Despite strong overall performance, challenges remain in region like Chittagong, where the models show lower precision and recall. Since no previous NER systems for Bangla regional dialects exist, our work represents a foundational step in addressing this gap. Future work will focus on improving model performance in underperforming regions and expanding the dataset to include more dialects, enhancing the development of dialect-aware NER systems.

命名实体识别方言识别低资源语言孟加拉语

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。