arXiv:2510.10776cs.CL2025-10

首个针对菲律宾希利盖农语的命名实体识别模型,解决低资源语言处理难题。

HiligayNER: A Baseline Named Entity Recognition Model for Hiligaynon

  • 基于8000+句标注语料,用mBERT和XLM-RoBERTa构建模型
  • 在所有实体类型上F1超80%,表现稳定可靠
  • 可跨语言迁移至宿务语和他加禄语,适合低资源语言研究

希利盖农语是菲律宾邦阿西兰、内格罗斯西部及苏克雷桑地区主要语言,但因缺乏标注语料与基准模型,长期被语言处理研究忽视。本研究提出HiligayNER,首个公开可用的希利盖农语命名实体识别基准模型。数据集包含超过8000条从新闻、社交媒体和文学文本中收集的标注句子。采用mBERT和XLM-RoBERTa两种Transformer模型在此语料上微调,构建HiligayNER版本。评估结果显示,两个模型在各类实体上的精确率、召回率和F1分数均超过80%。跨语言评估显示其在宿务语和他加禄语上具备良好迁移能力,表明HiligayNER在低资源多语言NLP中的广泛适用性。该工作旨在推动菲律宾弱势语言的技术发展,为区域语言处理研究提供基础支持。

原文摘要 · Abstract (English)

The language of Hiligaynon, spoken predominantly by the people of Panay Island, Negros Occidental, and Soccsksargen in the Philippines, remains underrepresented in language processing research due to the absence of annotated corpora and baseline models. This study introduces HiligayNER, the first publicly available baseline model for the task of Named Entity Recognition (NER) in Hiligaynon. The dataset used to build HiligayNER contains over 8,000 annotated sentences collected from publicly available news articles, social media posts, and literary texts. Two Transformer-based models, mBERT and XLM-RoBERTa, were fine-tuned on this collected corpus to build versions of HiligayNER. Evaluation results show strong performance, with both models achieving over 80% in precision, recall, and F1-score across entity types. Furthermore, cross-lingual evaluation with Cebuano and Tagalog demonstrates promising transferability, suggesting the broader applicability of HiligayNER for multilingual NLP in low-resource settings. This work aims to contribute to language technology development for underrepresented Philippine languages, specifically for Hiligaynon, and support future research in regional language processing.

命名实体识别低资源语言多语言迁移菲律宾语

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。